Tech
Build Once, Generate Endlessly: Putting the Synthetic Data Pipeline to the Test

Hyun Kim
Co-Founder & CEO | 2026/08/26 | 7 min read

In the first three installments of this series, we covered 5,000 actions, 50 spaces, and 10,000 object images. Each has value on its own, but fundamentally, they are raw materials. This final installment documents the first time those materials came together in a single pipeline.
The synthetic data work in Phase 2 was not about mass production. It was a proof of concept for the pipeline—a test of whether the three asset types could actually be combined inside a simulator to produce usable training data. This post covers both what we validated along the way and what remains to be solved.
What is synthetic data?
Training data generated by a simulator instead of captured in the real world. Because the simulator has access to the full state of every scene, ground-truth labels can be generated automatically.
Why Synthetic Data Changes the Cost Equation
Capturing a single one-to-five-minute sequence in the real world requires securing a location, actors, and a production crew. We experienced that cost structure firsthand in Phase 1.
By comparison, rendering a synthetic scene may take only seconds of GPU compute. Of course, the cost does not disappear. There is still an upfront cost to creating the space, action, and object assets that serve as the building blocks for synthesis.
The key difference is that these assets are reusable rather than consumed. Real-world filming incurs new costs for every scene, while asset creation is paid for once and then amortized across every scene generated from those assets. As the number of scenes grows, the cost per scene falls.

Conceptual diagram (relative comparison) — Asset creation costs are distributed across the number of scenes generated
The Pipeline: Bringing Three Asset Types Together in Isaac Sim
On top of the 3DGS environment from Part 2, we place objects built as physics-ready meshes from Part 3 and animate the digital humans created from the action assets in Part 1.
We then vary lighting, materials, camera viewpoints, and object placement through domain randomization to render multiple versions of the same scenario. The human model is designed to be body-shape-neutral, allowing body-type diversity to expand as additional assets become available.
Labels are generated automatically with each render. Bounding boxes and segmentation masks come by default, while the simulator can also provide information that is difficult to obtain from real-world capture—including depth, optical flow, pose, and 3D bounding boxes. In principle, anything the simulator knows about the scene can become a label.

Conceptual diagram — One scene, nine variations (not actual renders)
How We Validate Synthetic Data
The most important criterion for validating synthetic data quality is simple: does the rendered person interact correctly with the walls and floor?
A scene where a person's feet float above the ground or pass through the floor cannot be used as training data. The challenge is that automating this validation is extremely difficult.
Our current approach is iterative. We visualize the outputs and manually identify a type of error, write a test to detect that error, then visualize the results again to find the next failure mode. Through this loop, we have gradually built up layers of validation filters.
There are still limitations. Today, the spatial environment and animation data are extracted independently and then aligned during post-processing. In the next phase, we plan to explore joint optimization, where spatial and animation information are optimized together from the start so that both can be recovered simultaneously from captured video rather than reconciled afterward.
The Myth That More Photorealistic Is Always Better
There is a common assumption about synthetic data: the more photorealistic it looks, the more valuable it becomes.
Our experience building this pipeline suggests otherwise. The true value of a simulator is not how real it looks, but whether a model trained in simulation can successfully transfer to the real world.
Visual reconstruction, or real-to-sim, is a necessary foundation. But successful sim-to-real transfer requires much more: physical properties such as mass and friction, observation conditions such as sensor noise and latency, and deliberately constructed rare scenarios that are difficult to capture in the real world.
This is why we start with assets derived from reality—but do not treat visual realism as the end goal.
The Global Picture: Scale vs. Efficiency
Over the past year, China has established more than 40 robot training centers, taking a scale-driven approach to accumulating data through human and robot operating hours.
The path we validated takes a different approach. Instead of scaling primarily through physical facilities, it focuses on process efficiency: turning real-world data into reusable assets that can be reproduced in virtual environments.
The two approaches are not mutually exclusive. But for organizations that cannot compete purely on capital and workforce scale, the value of a reusable data production model becomes even greater.
Frequently Asked Questions
Q. Can robots be trained using synthetic data alone?
No. Real-world data provides the reference point, while synthetic data amplifies and complements it. Synthetic data is particularly valuable for improving edge-case coverage.
Q. What labels can be generated automatically?
Bounding boxes and segmentation masks are standard. The simulator can also output depth, optical flow, pose, 3D bounding boxes, and essentially any other scene information available within the simulation.
Q. What is domain randomization?
Domain randomization is a technique that deliberately varies environmental factors such as lighting, materials, camera viewpoints, and object placement during data generation. It helps prevent a model from overfitting to a specific environment and instead learn more fundamental features.
This concludes the series. The progression from raw data in Phase 1 → assetization in Phase 2 → validation of the production pipeline through a synthetic data PoC is what we mean by a Physical AI Data Factory. Full-scale production is planned for the next phase.
If your organization is exploring Physical AI data development or the use of synthetic data, talk to our team.
Superb AI is a vision intelligence company that transforms visual data from industrial environments into actionable intelligence for enterprises.
Related Posts

Tech
For a Robot, a Drawer Must Open, Not Just Be Seen: Building 10,000 Interactive Object Assets

Hyun Kim
Co-Founder & CEO | 22 min read

Tech
Not a Photo, but a Space You Can Walk Into: Reconstructing 50 Korean Homes with 3DGS

Hyun Kim
Co-Founder & CEO | 5 min read

Tech
How Robots Learn Human Behavior: Building 5,000 4D Behavior Data Assets

Hyun Kim
Co-Founder & CEO | 7 min read

About Superb AI
Superb AI is an enterprise-level training data platform that is reinventing the way ML teams manage and deliver training data within organizations. Launched in 2018, the Superb AI Suite provides a unique blend of automation, collaboration and plug-and-play modularity, helping teams drastically reduce the time it takes to prepare high quality training datasets. If you want to experience the transformation, sign up for free today.
