Insight

Synthetic Data for Defense AI Training: How to Build Training Data Where Real-World Data Is Scarce

Hyun Kim

Co-Founder & CEO | 2026/09/18 | 7 min read

Guide to Building Synthetic Data for Defense AI Training | Superb AI

Defense AI transformation is accelerating, beginning with infrastructure investment. In 2026, South Korea’s Ministry of National Defense allocated KRW 21.7 billion (approximately USD 15.7 million) to build GPU servers for a pilot of the Integrated Defense AI Data Center, while the government announced plans to establish an integrated defense AI data center with up to 50,000 GPUs by 2030 (ETNews). Work has also begun to integrate the data systems previously operated separately by each branch of the military into an “AI-Ready” architecture (News1).

Even with the compute infrastructure and data architecture in place, one question remains: What will the models be trained on? Training data for defense AI faces three constraints at once: security requirements, targets that are difficult or impossible to collect, and high data-development costs. This post explains how synthetic data can help overcome those constraints and reviews datasets Superb AI has built for defense and security applications.

Three Constraints on Training Data for Defense AI

First, security requirements. Video and image data held by the military is subject to restrictions on external transfer and third-party outsourcing depending on its security classification. Commercial cloud-based data development workflows can therefore be difficult to apply. What is needed is an environment where dataset development, training, and validation can all be completed within a closed network.

Second, scarcity. The scenes that surveillance and reconnaissance AI most needs to learn are often the hardest to collect in the real world. Actual threat scenarios, nighttime and adverse-weather conditions, and images of equipment that cannot be photographed domestically all offer very limited opportunities for collection. This creates a recurring cycle: without data, a model cannot be built, and without a model, there is little opportunity to accumulate more data.

Third, cost. Data with high demand in defense applications, such as satellite imagery, can be expensive even before annotation begins. In Superb AI’s Arirang satellite dataset project, high-resolution satellite images cost between KRW 600,000 and 800,000 per image, and each image contained more than 50 objects, making the labeling itself highly complex.

Where Synthetic Data Fills the Gap: Three Generation Methods

Synthetic data fills gaps in real-world data with artificially generated data. There are three main generation methods, which can be combined depending on the requirements and constraints of the task.

Conceptual illustration: The three synthetic data generation methods are not mutually exclusive and can be combined according to the constraints of the task

Real-data-based synthesis starts with available source data and composites the required objects onto it. AI for airport security X-ray screening is a representative example. Images of threat items that are difficult to place in actual baggage can be composited into normal baggage X-rays to create training data.

Read more → The Future of Airport Security—Synthetic Data for a World Where AI Learns on Its Own

Generative-model-based augmentation uses generative AI to change the lighting, weather, backgrounds, and other conditions in existing data to create additional variations. This is particularly useful for supplementing conditions that are difficult to collect, such as nighttime scenes or adverse weather.

Simulator-based synthesis builds environments and objects inside a 3D simulator and automatically generates labels along with the data. Because viewpoints, lighting, and object placement can be controlled programmatically, a single set of assets can continuously generate new variations. We covered a comparison of these methods and the engineering required to reduce the Sim-to-Real gap in a separate guide.

Read more → Building Superb AI’s Synthetic Data Pipeline with NVIDIA Isaac Sim

In Practice: What Superb AI Has Built for Defense and Security

Superb AI has more than 130 million industrial visual data samples and has built the following datasets for defense and security applications.

  • Satellite image object detection and classification: Detected and classified more than 500,000 objects across over 1,000 satellite images, achieving 99.9% semantic accuracy for object labels.
  • Crime and anomalous behavior detection: Achieved a 92.7% detection rate for anomalous behaviors such as falls and violence, and expanded the system into AI CCTV that automatically reports incidents after detection.
  • Missing-person search in inaccessible areas: Built an AI dataset for identifying the locations of missing persons in aerial imagery.

The synthetic data pipeline has already reached the test-operation stage. Superb AI is operating a pipeline that repeatedly generates training data inside a simulator using 5,000 human-action data samples, 50 indoor spaces, and 10,000 intelligent object assets as source material. The key is that once the source assets are created, their conditions can be changed to generate virtually unlimited variations.

Read more → Build Once, Generate Endlessly: Putting the Synthetic Data Pipeline to the Test

When to Use Synthetic Data—and When Not To

Synthetic data is particularly useful for three types of tasks: rare events with few opportunities for real-world collection, targets that cannot be captured at all, and large-scale data projects where source-data acquisition and labeling costs multiply quickly.

Conversely, if sufficient real-world data is already available, there is no reason to begin with synthetic data. Even when synthetic data is used for training, model performance should be validated on a real-world evaluation set. Quality also depends on curation: rather than using every generated sample, the data that contributes meaningfully to training should be selected.

Security requirements can be addressed through the deployment architecture. Superb AI supports on-premises deployment, allowing dataset development and model training to take place within a closed network without transferring source data outside the organization. Superb AI has also continued working in the defense sector, including participation in the Korea Army International Defense Industry Exhibition (KADEX) in 2024.

Frequently Asked Questions

Q. Our security requirements prohibit us from sending data to an external provider. Can we still build the dataset?

Yes. With an on-premises deployment, the platform can be installed within the customer’s closed network so that the source data does not leave the environment. Security procedures for implementation personnel can also be designed together at the start of the project.

Q. Can an AI model be trained using only synthetic data?

We do not recommend it. Training should combine synthetic and real-world data, while performance validation should always use a real-world evaluation set. Synthetic data is not a replacement for real-world data; it is a way to fill the gaps where real data is insufficient.

Q. What types of projects should consider synthetic data first?

Start with projects where real-world data collection is blocked or severely constrained. Typical examples include rare-event detection, recognition of targets that cannot be photographed, and augmentation of nighttime or adverse-weather conditions. For tasks where real-world data can be collected, compare the cost of building the dataset with real data before deciding.

Q. Won’t models trained on synthetic data perform worse in real-world environments?

The Sim-to-Real gap is a real issue. Standard practice is to reduce that gap by increasing variation through domain randomization and repeatedly validating performance against a real-world evaluation set.

💬 Having trouble securing enough training data for your AI project? Tell us about your use case and data conditions—including security classification and whether real-world collection is possible—and we’ll start by assessing whether synthetic data is a suitable option.

Superb AI is a Vision Intelligence company that transforms visual data from industrial environments into actionable intelligence for enterprises. With more than 130 million industrial visual data samples, Superb AI builds AI training datasets and synthetic data pipelines designed around customers’ security requirements.

About Superb AI

Superb AI is an enterprise-level training data platform that is reinventing the way ML teams manage and deliver training data within organizations. Launched in 2018, the Superb AI Suite provides a unique blend of automation, collaboration and plug-and-play modularity, helping teams drastically reduce the time it takes to prepare high quality training datasets. If you want to experience the transformation, sign up for free today.