Case Study

[Customer Success Story] How AgTech Company Metafarmers Built Crop Recognition Datasets for Harvesting Robot Vision

Hyun Kim

Co-Founder & CEO | 2026/09/11 | 5 min read

How AgTech Company Metafarmers Built Crop Recognition Datasets for Harvesting Robot Vision

Efforts to automate crop harvesting—work that has traditionally relied heavily on manual labor—are rapidly growing into an industry of their own. MarketsandMarkets projects the global agricultural robots market to grow from USD 17.73 billion in 2025 to USD 56.26 billion by 2030, representing a CAGR of 26.0% (MarketsandMarkets, Agricultural Robots Market Report). As labor shortages in rural areas become more structural, demand is rising to automate labor-intensive tasks such as harvesting.

Before a harvesting robot can work in the field, it first needs to recognize the crops around it. The perception model that determines what should be harvested is a fundamental part of the robot—and the performance of that model depends on its training data.

Metafarmers, an AgTech company developing AI-powered robots for automated crop harvesting, came to Superb AI with this challenge. The company needed training data for its crop recognition models and wanted to know whether the labeling work could be handled externally. Superb AI combined Superb Platform with its data services to build the required crop recognition datasets, establishing a repeatable data pipeline for Metafarmers.

Challenge: Robotics Talent Tied Up in Labeling

The harvesting robot’s perception model required two types of data. One dataset was needed to recognize young crops at the seedling stage, while another was needed to identify harvest targets among fully grown crops. Because crop appearance changes across growth stages, each stage required a separate dataset.

The challenge was that crop images are particularly tricky to label. Leaves and stems overlap, and a single object may appear to split into multiple branches. Without clearly defined annotation criteria for these ambiguous cases, consistency across the dataset can quickly break down.

Until then, Metafarmers had labeled the data it needed internally. But as data requirements grew, having the same engineers responsible for developing robots and models spend their time on annotation became increasingly difficult to sustain. To keep robot development and dataset production from competing for the same resources, the company needed an external labeling pipeline it could rely on.

Solution: Two Annotation Tasks Tailored to Different Recognition Needs

Superb AI designed the project as two separate annotation tasks. Seedling recognition data was labeled with bounding boxes, while harvest-target data was processed using instance segmentation, which captures the precise outline of each object. The annotation method was selected to match the level of precision required for each recognition task.

Figure 1. Two annotation methods tailored to different crop growth stages (conceptual illustration; not actual customer data)

Before annotation began, the teams aligned on the criteria. Metafarmers already had internally labeled data and an existing labeling guide, so Superb AI built on those materials rather than starting from scratch. During kickoff, the teams reviewed the annotation standards together and documented ambiguous cases specific to crop imagery before production began.

Figure 2. Dataset development workflow—from building on the customer’s existing guidelines to a repeatable production process (reconstructed)

Benefit: Keeping the Team Focused on Robot Development

The project delivered two key outcomes for Metafarmers. The first was the crop recognition data needed to train its robotic perception models. The second was an external data pipeline that could be run repeatedly whenever new data needed to be processed.

By moving annotation outside the internal development team, Metafarmers created a structure in which its engineers could stay focused on robot and model development.

Figure 3. How the training data acquisition process changed (conceptual illustration) 

For Superb AI, this case demonstrates how data services can support robotics companies beyond a one-off labeling project. The process starts by building on guidelines the customer already has, adapting the project scope and annotation approach to the characteristics of the data, and applying the same standards again as new tasks and datasets emerge.

For robots to perceive, make decisions, and act in the physical world, they first need data that teaches them what to recognize. Superb AI helps build that data foundation.

Frequently Asked Questions

Q. Can labeling quality be maintained for irregularly shaped objects such as crops?

Yes, if the criteria are defined first. Ambiguous cases—such as overlapping leaves and stems—should be addressed in the annotation guidelines during kickoff, and both labeling and quality review should follow those standards. In this project, Superb AI used the customer’s existing guidelines as the starting point and refined the criteria together before annotation began.

Q. We already have data that we have been labeling internally. Can Superb AI continue from it?

Yes. In this case, Superb AI built on the customer’s existing labeled data and annotation guidelines, aligned on the criteria before production began, and allowed the customer to perform the final review directly in Superb Platform.

If you have any questions during review, please feel free to reach out to your Superb AI contact at any time.

About Superb AI

Superb AI is an enterprise-level training data platform that is reinventing the way ML teams manage and deliver training data within organizations. Launched in 2018, the Superb AI Suite provides a unique blend of automation, collaboration and plug-and-play modularity, helping teams drastically reduce the time it takes to prepare high quality training datasets. If you want to experience the transformation, sign up for free today.