Case Study
[Customer Success Story] How a Physical-Space Video Analytics Company Built Pedestrian Detection and Tracking Datasets with De-Identified Data

Hyun Kim
Co-Founder & CEO | 2026/08/28 | 7 min read
![[Customer Success Story] How a Physical-Space Video Analytics Company Built Pedestrian Detection and Tracking Datasets with De-Identified Data](https://cdn.sanity.io/images/31qskqlc/production/2590bf183ae90b8945e3f324bb55d15c798cc65e-3200x1800.png?fit=max&auto=format)
The video analytics market, which uses AI to extract meaningful insights from video, is growing rapidly. Mordor Intelligence projects the global video analytics market to grow from USD 12.39 billion in 2025 to USD 33.74 billion by 2030 (Mordor Intelligence, Video Analytics Market Report). As the market expands, so does demand for training data. But when people appear in video footage, that data inevitably comes up against privacy regulations.
In Korea, this tension is reflected in the regulatory framework. The use of original video footage for AI training is restricted in principle. Even when the Ministry of Science and ICT permitted local governments to use original CCTV footage for AI training through a special demonstration exemption under the ICT Regulatory Sandbox in December 2025, the approval came with conditions: compliance with the Personal Information Protection Commission’s safety standards for the use of original video data and use of a designated Data Safe Zone (Edaily, December 23, 2025). The fact that an exceptional regulatory process is required illustrates just how high the barrier to using original footage remains. In practice, de-identification is therefore the more viable path for most companies.
A vision AI company specializing in video analytics for physical spaces faced the same challenge. AI models designed to understand human movement ultimately need to learn from data that contains people—but the videos and images used for training can also contain personally identifiable information, including faces.
To improve a model that recognizes the location and movement of people in video, the company needed a pedestrian detection dataset. There was one key requirement: the images used for labeling had to be de-identified rather than original images. Superb AI designed a data labeling pipeline around this requirement, first building a detection dataset and later expanding the project to include a video-based tracking dataset.

Challenge: Building Training Data While Protecting Personal Information
What sets pedestrian dataset development apart from conventional object detection labeling is the sensitivity of the underlying data. When people are the subjects of an image, the data is subject to Korea’s Personal Information Protection Act, making it impractical to provide original images directly to external annotators.
When the company first approached Superb AI, its primary questions were not about technology, but compliance: what type and level of de-identification should be applied, and could labeling quality still be maintained on de-identified images?
Stronger de-identification reduces privacy risk, but it can also remove visual information that annotators need to work accurately. The right balance between privacy protection and labeling quality therefore had to be defined before the project began.
There was also a practical resource constraint. It was difficult for the company’s internal researchers to handle labeling alongside model development. Having researchers spend their time annotating thousands of images was not a sustainable use of specialized resources.
Solution: A Labeling Pipeline Designed Around De-Identification Requirements
During kickoff, Superb AI reviewed the project requirements with the customer and established the de-identification criteria and labeling guidelines before annotation began. The source data consisted of de-identified pedestrian images, with pedestrians labeled using bounding boxes and additional attributes recorded for the person class.
A dedicated project manager and technical support engineer were assigned to the project because the engagement involved more than labeling alone. Two technical workflows ran in parallel.
First, the customer’s existing pre-labeling results were imported into Superb Platform and used as the starting point for annotation. Second, the completed labels were converted into the format required by the customer’s training pipeline before delivery.
Rather than discarding existing work and starting from scratch, Superb AI built on the data already available and focused annotation efforts only where additional work was needed.

Expansion: From Image Detection to Video Tracking
As the detection dataset neared completion, the customer requested the next stage of data development: a tracking dataset that went beyond identifying people in individual images to maintaining the identity of the same person across consecutive video frames. A second engagement to build the tracking dataset followed within the same year.
Moving from detection to tracking changes the unit of data from individual images to video sequences and introduces an additional requirement: assigning and maintaining a consistent ID for the same person across frames. This increases both labeling complexity and quality-control requirements.
However, because the de-identification criteria, labeling guidelines, and project team had already been established during the first engagement, the tracking project could begin without a separate setup period.
Benefit: Keeping Researchers Focused on Models—and Data in the Pipeline
The project delivered two key outcomes for the customer.
The first was a training dataset built while meeting applicable privacy requirements. Because the level of de-identification was agreed upon and documented in advance, the project also preserved documentation supporting the compliant use of the data alongside the final deliverables.
The second was research time. By moving annotation and format conversion into an external data pipeline, the company was able to keep its internal researchers focused on model development rather than large-scale manual annotation.

For Superb AI, this case demonstrates how data labeling can become part of a customer’s broader data pipeline rather than remain a one-off outsourcing task. Existing pre-labeling results can be carried forward, output formats can be adapted to the customer’s training environment, and the same standards can scale as the task evolves from detection to tracking.
For Superb AI’s vision of delivering Vision Intelligence for industry, building the right data foundation is where that process begins.
Frequently Asked Questions
Q. Can labeling quality be maintained on de-identified images?
It depends on the level of de-identification. That is why Superb AI defines the type and degree of de-identification together with the labeling guidelines during project kickoff. Agreeing in advance on a level that protects personal information while preserving the visual information needed for annotation is essential to maintaining quality.
Q. We already have data that we labeled internally. Can Superb AI continue from it?
Yes. In this case, the customer’s existing pre-labeling results were imported into Superb Platform and used as the starting point. The completed annotations were then converted into the format required by the customer’s training pipeline.
Q. Does Superb AI support short-term projects involving only a few thousand images?
Yes. The first engagement in this case also involved several thousand images. Superb AI assembles a dedicated project manager and annotation team based on the project’s scale and timeline, and the same standards can be carried forward as the scope expands.
Related Posts
![[Customer Success Story] Generative AI Starts with a Blueprint: AI ISP Consulting for a Public-Sector Institution](https://cdn.sanity.io/images/31qskqlc/production/17a72bed2d9fcfe14f67c854211991f6dbd795a7-3200x1800.png?fit=max&auto=format)
Case Study
[Customer Success Story] Generative AI Starts with a Blueprint: AI ISP Consulting for a Public-Sector Institution

Hyun Kim
Co-Founder & CEO | 6 min read
![[Customer Success Story] Five Years of Advancing Perception AI: How an Autonomous Robot Service Company Built a Multi-Sensor Data Labeling Pipeline](https://cdn.sanity.io/images/31qskqlc/production/6967e3a8b41f177b722361c7fcad9711ce231cea-3200x1800.png?fit=max&auto=format)
Case Study
[Customer Success Story] Five Years of Advancing Perception AI: How an Autonomous Robot Service Company Built a Multi-Sensor Data Labeling Pipeline

Hyun Kim
Co-Founder & CEO | 7 min read
![[Customer Success Story] From Component Identification to Defect Detection: How a Manufacturer of Industrial Automation Components Built a PCB Labeling Workflow](https://cdn.sanity.io/images/31qskqlc/production/40834bdafb10540bb7c5f90aa6731a59976ec68a-2560x1434.png?fit=max&auto=format)
Case Study
[Customer Success Story] From Component Identification to Defect Detection: How a Manufacturer of Industrial Automation Components Built a PCB Labeling Workflow

Hyun Kim
Co-Founder & CEO | 10 min read

About Superb AI
Superb AI is an enterprise-level training data platform that is reinventing the way ML teams manage and deliver training data within organizations. Launched in 2018, the Superb AI Suite provides a unique blend of automation, collaboration and plug-and-play modularity, helping teams drastically reduce the time it takes to prepare high quality training datasets. If you want to experience the transformation, sign up for free today.
