When a drone sees and decides: what AI can do on the battlefield today and where humans must remain involved
How a drone recognizes its surroundings, why it runs a small YOLO model instead of a giant language model, and why context is the hardest part to understand. Technology, autonomy, law and Czech research without unnecessary scaremongering.

The title illustration refers to the fictional drone Baron from my book Syndikát. It does not depict the actual product or a real military operation.
CNN Prima NEWS · September 5, 2026 at 10:10 a.m.
Today I spoke in a live interview about what artificial intelligence means inside a drone, where assistance ends and autonomy begins, why human decision-making can be too slow on the modern battlefield, and what AI, despite its rapid development, cannot reliably judge. This article expands those few minutes on television with technical context, data, law and Czech research.
When one says "artificial intelligence drone", one easily imagines a single electronic brain that sees, thinks and makes decisions. The reality is more sobering. Inside, a camera, several small specialized models, navigation algorithms and a classic autopilot usually work together.
Baron from the opening image belongs to fiction. The technologies described in the following text are real, and I deliberately separate them from the story.
Each part solves a different problem. The camera delivers the image. The computer vision model marks the objects. Navigation will estimate the location. The planner selects a safe route. The autopilot will translate the plan into motor commands.
Recognizing the tank in the image is a technical problem. Deciding whether it is a legitimate military objective is a question of context, law and human responsibility.
It is this difference that I consider to be the most important part of the whole debate.
What really happens inside the drone
A simplified processing chain looks like this:
| Step | What's going on | Typical Tools |
|---|---|---|
| Perception | A camera or thermal camera provides an image and the system searches for objects in it. | YOLO, Image Segmentation, Object Tracking |
| Positioning | The drone estimates its position and orientation. | gyroscope, accelerometer, GPS/GNSS, visual odometry, map matching |
| Planning | The system compares the route, obstacles, prohibited zones and energy status. | navigation algorithms and mission rules |
| Flight Control | The autopilot converts the required direction into the speeds of the individual motors. | control loops and flight controller |
So AI is not synonymous with autonomy. The drone can only use AI to recognize obstacles, while every important step is approved by a human. Conversely, some automatic functions can be built on classical algorithms without a neural network.
In addition, the model usually does not learn everything from scratch on the fly. Computationally intensive training takes place earlier on large datasets. The drone then mainly runs inference, i.e. the quick application of the learned model to a new image.
YOLO is not a small ChatGPT. It provides fast visual perception
In the interview, I mentioned the YOLO family of models, from English You Only Look Once. It is not a language model, but a computer vision model for object detection. It will usually return the rectangles, class names and confidence level from the image. For example, it can say "there is a vehicle in this part of the image, 87% confidence".
YOLO does not know if the vehicle belongs to its own unit, if it is disabled, if it surrenders, what is behind it, or if the intervention would be reasonable under the law. These are all additional layers of decision making.
Each generation of YOLO has several sizes from n as nano to x as extra large. The following numbers are from the Ultralytics YOLOv8 documentation and the official YOLOv12 repository. The last column is my approximate recalculation of the weights themselves at two bytes per parameter in FP16 format.
| Variant | YOLOv8 | YOLOv12 | Approximate weight size in FP16 |
|---|---|---|---|
n nano | 3.2 million parameters | 2.5 million parameters | approximately 5 to 7 MB |
s small | 11.2 million | 9.1 million | 18 to 23 MB |
m medium | 25.9 million | 19.6 million | 39 to 52 MB |
l large | 43.7 million | 26.5 million | 53 to 87 MB |
x extra large | 68.2 million | 59.3 million | 119 to 136 MB |
In comparison, YOLOv12n has 2.5 million parameters. A common language model referred to as 7B has seven billion. Thus, according to the number of parameters, YOLOv12n is about 2,800 times smaller. The parameter numbers of the largest GPT or Gemini closed models are not public, so it makes no sense to pass them off as a verified number.
However, the size of the weights file is not the same as the memory consumption during operation. The drone also needs space for intermediate layer results, image buffers, and the runtime environment. It depends on the resolution, the accuracy of the calculation, the amount of images, the chip and the specific engine. An honest number will therefore only be created by measuring on the target device.
YOLOv8 is mainly a convolutional network. YOLOv12 is a newer branch of research that relies more heavily on the attention mechanism. A higher generation number alone does not mean the best choice for a particular drone. In practice, accuracy, latency, power consumption, tool stability and whether the model has passed tests on real hardware are decisive.
Why intelligence must be right on board
You can't simply put a server graphics card with a large cooling and power supply in a small flying device. Every gram and every watt translates into flight time. In turn, sending each frame to the cloud adds delay and assumes a reliable connection that may not exist in a crisis environment.
That is why edge AI is used, i.e. calculation directly on the device. It brings three fundamental advantages:
- fast response without data path to a remote server,
- less dependence on connection and cloud,
- the ability to work even when communication or satellite navigation is interrupted.
The price for this is clear: low power, low memory and limited processing power. The developer therefore shrinks, quantizes and optimizes the model for a specific accelerator. Success is measured not only by accuracy in the lab, but also by frames per second, power consumption, temperature and signal loss behavior.
DARPA REMA Program, for example, is researching autonomous subsystems that can help existing drones complete a mission even when communications are disrupted by electronic warfare. This demonstrates well why edge computing is of military interest. At the same time, it is precisely the loss of connection that increases the demands for predetermined restrictions and the safe termination of the mission.
A gyroscope, an accelerometer and a map on your knee
The image alone is not enough. The drone needs to know how it is turned and where it has moved.
The Gyroscope measures the angular velocity, i.e. how fast the device rotates. Accelerometer measures acceleration in individual axes and gravity is also reflected in its data. Together they form the basis of the inertial measurement unit, IMU for short.
If we just add up the data, small errors will gradually add up and the position estimate will start to slip. Therefore, the IMU is combined with a camera, satellite navigation or a map.
I used the analogy of a World War I pilot on TV. With a paper map on his knee, he looks out, finds a river, a railroad, and a town, and corrects his estimate of position based on their shape. Map matching does something similar by machine: it compares the observed features of the landscape with a pre-prepared digital map.
With modern systems, visual-inertial odometry or SLAM is often added, i.e. simultaneously building a map and estimating its own position. A study of an autonomous racing drone Swift in Nature showed what a combination of cameras, IMUs and on-board computing can do in a controlled environment. However, the race track is still incomparably clearer than the battlefield with smoke, interference, false targets and civilians.
Three levels: human in, on and out of the loop
The terms human in the loop, human on the loop and human out of the loop are useful shorthand, not a uniform legal norm. In practice, it is more important to ask specific questions: Who set the mission? Who determined the target profile? Who chooses a specific object? Can one stop the action in time?
| Mode | System role | Human role |
|---|---|---|
| Human in the loop | AI recommends, tracks or navigates. | A human authorizes a specific use of force. |
| Human on the loop | The system acts in a predefined space and time. | A person supervises and has a real possibility to intervene or end the action. |
| Human out of the loop | Once activated, the system selects and attacks specific objects without further human intervention. | A human sets the rules beforehand but does not decide on each individual attack. |
There is a wide spectrum between a manually piloted drone and full autonomy. AI can only stabilize the flight, suggest a route, draw attention to an object, or complete the last section of guidance after the target has been selected by a human. Therefore, the phrase "the drone decided on its own" means almost nothing without further explanation.
Also, the word target has at least three meanings:
- strategic purpose, for example to protect territory,
- the profile of the object in the mission, for example, to search for a certain type of equipment in a specified zone,
- the specific object against which the force is to be used.
The further down this list human decision-making is removed, the more important are the constraints of space, time, object types, behavior under uncertainty, and the possibility of safe interruption.
Hyperwar: when the decision cycle is shortened to machine speed
The term hyperwar was used by John R. Allen and Amir Husain for a conflict in which automated decision-making and concurrent actions occur at a speed that is difficult for humans to keep up with (US Naval Institute, 2017).
It's tempting to say that the faster side automatically wins. But speed can also reduce the time for review, doubt and de-escalation. A single sensor error then develops not over hours, but over seconds and on a large scale.
Therefore, meaningful human control cannot mean a human just watching dozens of automatic decisions on a screen. It requires:
- enough information about the situation and uncertainty of the model,
- the time in which the decision can actually be changed,
- a manageable number of simultaneous events,
- a clear scope of permitted activity,
- possibility to stop the system and find out why it acted.
A person who has a fraction of a second to react and does not see the reasons for the decision is part of the process in name only.
The hardest thing is not to recognize the tank. The hardest part is understanding the scene
In a clean photograph, it can be relatively easy to recognize a known type of equipment. But the battlefield is an environment outside of textbook data. Objects are covered, damaged, differently lit, viewed from an unusual angle, or intentionally camouflaged. The sensor may be damaged and the adversary is trying to confuse the system.
NIST in its Taxonomy of Adversarial Machine Learning describes attacks on the data as well as on the model's decision making itself, and notes that there is no one-size-fits-all defense.
Even more difficult is the meaning of the scene. The detector can see the vehicle, but does not know:
- who does it belong to and whether it has been captured,
- whether it is operable or abandoned,
- whether its crew is surrendering,
- whether there are civilians, medical personnel or a protected object nearby,
- what military benefit and what collateral damage would the intervention bring.
This is the difference between detection, identification, situational understanding and legal decision. Marketing sometimes merges them into a single word "intelligence". In fact, there is a gulf between them.
There is also a human risk called automation bias. The operator can trust the machine precisely because the result looks accurate: the frame, the class name, the map and the percentage of certainty. However, the 92% number says something about the model and its data, not whether the intervention is correct and legal.
What is already happening in practice
The pace of development is high. Ukraine's Ministry of Defense reported in April 2026 that over 200 companies are working on drones with AI elements in the country, the Brave1 platform records more than 300 related developments, and over 70 AI or computer vision solutions are used on the frontline (Ministry of Defense of Ukraine).
These numbers are an official statement from one warring party, not an independent audit of effectiveness. However, they show the speed of experimentation well. At the same time, the same ministry wrote in May that it does not seek 100% autonomous combat systems and that the final decision should remain under human control (Defense AI Center A1).
Law does not wait for a new type of machine
There is not yet a separate global binding convention for autonomous weapon systems. It does not mean a legal vacuum. According to the International Committee of the Red Cross, any attack is subject to international humanitarian law, regardless of the technology used.
Three duties are key to a particular attack:
- Distinction: distinguish military targets from civilians and civilian objects.
- Proportionality: avoid attacks expected to cause civilian harm excessive in relation to the concrete and direct military advantage anticipated.
- Precautions: take practicable steps to verify the target and minimize civilian damage, including changing or canceling the attack if the situation changes.
Responsibility does not disappear inside the algorithm. States, commanders, operators, developers, and procurers have different roles, but the phrase "AI decided" is not a legal or moral excuse.
ICRC in June 2026 again called for a binding international instrument that would ban unpredictable autonomous weapons and anti-personnel systems and severely limit others. Discussions of the Expert Group under the Convention on Certain Conventional Weapons continued in 2026. The CCW Review Conference is scheduled for November 2026.
Czechs have a surprisingly strong school in autonomy
The Czech Republic is not the main producer of cheap mass-produced FPV drones. However, it has a very good research base in robotics, swarms, computer vision and non-GPS navigation. It is these abilities that will be important not only in defense, but also in rescue, infrastructure control and work in dangerous environments.
ČVUT FEL, Professor Martin Saska's Multi-robot Systems Group develops safe flying robots, swarm coordination and navigation without GNSS or without continuous human connection. The team applies its results in the protection of critical infrastructure in European projects (ČVUT FEL, 2026).
The Drone Research Center at the BUT in Brno reports that it has been conducting, since 2020, drone swarm research projects for the Ministry of Defense and the Army of the Czech Republic. It combines control, electronics, sensors, AI and GNSS-free operation (VUT DRC).
The University of Defense and BUT jointly developed the ROJ project, in which aerial drones and ground robots work together during reconnaissance. AI divides tasks, evaluates the image and alerts the operator. It is important to add that the publicly described purpose is reconnaissance and information gathering, not autonomous decision making to use lethal force (Defense University).
This is a Czech advantage that we can rightly be proud of: not one miracle box, but a long-term research tradition in control systems, robotics and machine cooperation.
Better than automating war is to help prevent it
I consider the development of defense technologies to be legitimate and sometimes necessary. But the use of force must be a last resort. Technical ability alone does not answer the question of whether a particular deployment is correct.
Therefore, for me, the topic of drones also includes prevention, societal resilience and education. In the project DigiVoják a DigiVojákyně we teach pupils and students to recognize disinformation and manipulation, understand the role of AI, prepare for crisis situations such as blackouts, and think about security as part of responsible citizenship.
It's not about war scares. It is about the ability not to panic, verify information, cooperate and understand technologies that will influence the free decisions of the entire society.
My conclusion
AI can give the drone quick vision, more accurate stabilization, navigation without a constant connection and the ability to react to obstacles. However, it has no human understanding of the situation, no conscience, and no legal responsibility.
Therefore, the most important question is not: Can a machine make a decision faster than a human? It often already can.
The real question is: Which decisions are we even allowed to entrust to it, how do we recognize its uncertainty and who will be responsible if it makes a mistake?
Speed is a technical advantage. Responsibility is a human duty.
Note: The article is educational. It does not describe weapon design, tactical guidance, or target selection and engagement procedures. The links lead primarily to official documentation, professional publications and public information of universities and institutions.
Resources and further reading
- Ultralytics: YOLOv8 and Official YOLOv12 Repository
- Nature: Champion-level drone racing using deep reinforcement learning
- NIST: Adversarial Machine Learning, Taxonomy of Attacks and Defenses
- ICRC: AI in the military domain
- ICRC: Drones and international humanitarian law
- US Naval Institute: On Hyperwar
- ČVUT FEL: European drones of the new generation
- VUT Drone Research Center and Defense University: ROJ project
- Syndikát: the book that the Baron drone theme comes from