On one side, robotics companies create orders to fuel fundraising and listings; on the other side, data is endlessly discounted and subcontracted through layer after layer. Caught in the middle, data collection firms criticize this model while simultaneously profiting from it.
Old Tai (a pseudonym) has been in the embodied AI data collection business for a year or two, and lately he has started to feel lost. His company has grown from just a few people to several dozen, managing over a thousand collectors across multiple small third- and fourth-tier counties nationwide through housekeeping agencies, staffing companies, and crowdsourcing channels. If orders go smoothly, he projects revenue approaching 100 million yuan this year. But now he finds it hard to say what this business will look like next year.
Over the past two years, robot data collection centers have sprung up. They build household and factory scenarios where collectors remotely operate robots to repeatedly perform tasks like tightening screws, picking up cups, and folding clothes, then record the robots' motion data. This kind of real-machine data carries high equipment and labor costs, and industry insiders say the market price is roughly 500 to 1,000 yuan per hour. Currently, robots have not yet entered factories and homes at scale, and the lack of data is one of the core challenges for application, which is why data collection and trading is an important business in the embodied AI industry.
This year, more companies are shifting toward lower-cost bodyless collection. The collection process no longer relies on robots; instead, ordinary people wear cameras, grippers, and other devices to complete tasks like organizing, carrying, and cleaning, and this data is then used for model pre-training. The advantage of this approach is large volume and low cost. Data sales mainly target robot body manufacturers and model companies like Unitree and AgiBot. Data is used both to train models and to generate procurement orders.
Several robotics companies approached Old Tai, hoping both sides could make orders by purchasing from each other, using hardware and data to create transaction revenue. He also once participated in a deal with no physical delivery: both parties signed contracts separately, one buying components and the other buying data, with funds circulating once around and each side keeping the revenue. This year, more than twenty robotics companies plan to list in Hong Kong, and according to multiple media reports, embodied AI listing reviews are tightening, with regulators paying more attention to whether order revenue is real and sustainable. Even for a leading embodied AI company like Unitree Technology that has already achieved profitability, its large-scale application is still at an early stage. According to its financial report, in the first half of 2026, Unitree's revenue was 1.152 billion yuan, with net profit attributable to shareholders of 274 million yuan. From the perspective of application scenarios, in the first three quarters of 2025, among Unitree's humanoid robot revenue, research and education scenarios contributed 73.6%, commercial and consumer scenarios accounted for 17.39%, and industrial applications accounted for 9.01%. Another leading robotics company, UBTECH, has been listed on the Hong Kong Stock Exchange for nearly three years and is still not profitable. According to its earnings announcement, in the first half of 2026, the company achieved operating revenue of 1.269 billion yuan and a net loss attributable to shareholders of 311 million yuan.
Over the past year, Old Tai has watched monocular data drop from over 200 yuan per hour to just over 100 yuan. Data priced at just a few yuan per hour has even appeared on the market. This data passes through multiple layers of subcontracting, with unclear origins and no way to know how many times it has been sold. On one side, robotics companies create orders to fuel fundraising and listings; on the other side, data is endlessly discounted and subcontracted through layer after layer. Caught in the middle, Old Tai criticizes this model while simultaneously profiting from it.
Trading and fundraising
Over the past few months, some robotics companies sprinting toward listing have approached Old Tai, hoping both sides could purchase from each other, using hardware and data to create transaction revenue for each other. Old Tai once "ran a set of books" for a robotics company, signing a contract worth about several million yuan. One side purchased hardware, the other purchased data, and neither completed physical delivery—only account transactions were made. He also knows there are capital and tax risks involved. When asked why he participated, Old Tai said it was just to "make a friend." After completing this transaction, someone else approached Old Tai, hoping to borrow his company as an intermediary to "do a favor." Under their plan, one company would first transfer money into Old Tai's company account, and then he would hand part of the funds to designated people in cash. Transaction records would remain between companies, while cash would go to individuals. Old Tai did not agree, saying he would "think about it." More complex transactions would involve component companies, suppliers, and subsidiaries. Funds would pass through three or four entities before returning to the starting point, with each company obtaining an order. Whether equipment is mass-produced or whether data is used for training is very difficult for outsiders to judge.
A peer told Old Tai about a data company that completed multiple rounds of financing in a short period, but through crowdsourced data collection, "legally extracted" investors' money. Equipment prices can be inflated in contracts. According to him, crowdsourced data has no unified pricing, and investors find it very difficult to see the true costs. For example, a set of equipment with an actual price of about 600,000 yuan could be signed into a contract at a listed price of 1 million yuan. The buyer pays 1 million yuan first, and the seller then returns 400,000 yuan through other means. If the contracted business is also an affiliated company, the financing funds can be transferred out in the form of procurement expenses.
In 2026, China's embodied AI sector is experiencing a dense wave of financing. According to data from third-party institution IT Juzi, in the first half of the year alone, total financing in China's embodied AI sector reached 93.5 billion yuan, a fivefold increase compared to the first half of 2025; the number of financing events reached 322, a year-on-year increase of 137%. From Old Tai's perspective, investors rarely ask how data is collected or whether it can be used for training. They care more about the number of devices, order amounts, and operating revenue. Data is hard to price, while hardware and contracts are easier to write into financing materials. The bundled play of investment and orders is very common in the industry. While the invested company receives funds, it also takes on a procurement task. Old Tai said that an investment-to-procurement ratio of 1:0.5 or 1:0.7 is relatively reasonable, but there are also ratios of 1:2 or 1:3, which he described as "quite ruthless."
Local data collection factories have also been pulled into this trading scheme. A common practice is for robotics companies to establish joint ventures with local investors to carry out embodied AI-related data collection business. The joint venture purchases a batch of robots from the robotics company for data collection, the robotics company records equipment shipments, and then promises to buy back a certain scale of data in the coming years. The problem is that data prices are falling. Old Tai understands that data collection factories have already invested in equipment and personnel costs, but robotics companies can delay buybacks or renegotiate prices downward. Different robots also use different data formats, making it very difficult to sell this data to other customers. He has encountered some data collection factories whose cost of producing one hour of data reaches 500 to 1,000 yuan, yet they ultimately sell it for just a few hundred yuan and still cannot find buyers. Robots and equipment remain on site, and the data cannot be sold either.
Data prices keep falling
A body manufacturer client once proposed wanting hundreds of thousands of hours of data within the month. Old Tai found this difficult, saying it was not a task that could be completed by just adding a few people on short notice. Data collection has become a business of organizing labor. Orders flow from body manufacturers all the way down to county towns. After receiving demand from body manufacturers or model companies, Old Tai distributes collection devices to housekeeping companies, staffing companies, and crowdsourcing teams. These contractors then subcontract further downward, ultimately reaching ordinary people. They carry collection devices during their daily work and produce data as a side job, mostly in third- and fourth-tier cities and county towns. The devices are not complicated either. Phones, GoPros, and headbands can all record first-person video. Phones and GoPros typically collect monocular data. Binocular devices use two lenses simultaneously to record more spatial information. As demand increased, more data companies, staffing companies, and intermediaries entered. Old Tai says data that could sell for over 200 yuan per hour last year can only sell for just over 100 yuan this year. As procurement volumes increased from a few thousand hours to tens of thousands of hours, clients continued to push prices down. To make money, they can only keep finding people further down. The industry initially sold data externally at 120 yuan per hour, or after passing through five or six layers of intermediaries, collection eventually reached rural areas where only 5 yuan per hour remained. Data flows toward cheaper places, with each intermediary taking a slice of profit, and devices are distributed layer by layer through multiple tiers of suppliers. Data flows toward cheaper labor. Some go to rural areas, some to Southeast Asia, and others seek out more niche regions. As long as local people can be organized and there is a device that can record video, collection prices can continue to drop. Old Tai pays at least 30 to 35 yuan per hour to find storefronts in county towns; Southeast Asian workers have lower wages, but after labor agency commissions, costs are similar to domestic ones. Yet market prices have already fallen below this line. Data prices range from several hundred yuan to single digits, with some selling data at 5 yuan or even 3 yuan per hour. Most of these single-digit quotes are not prices from normally organized collection. Old Tai says that just calculating collector wages, the cost of one hour of data is more than single digits. Low-priced data usually comes from subcontracting chains, and may also be data that was privately retained and resold. The industry calls this kind of data with unclear origins and authorization relationships "black data."
Data ownership is very difficult to trace. When Old Tai signs contracts with clients, resale by clients is generally prohibited. But this constraint is very difficult to enforce. "You sell it once, he can sell it five times," he said. Data is a difficult problem for the robotics industry. Unlike large language models, where the internet has massive amounts of data to collect, robots have basically had no real-world data before. The industry hopes to first expand data scale in the hope of replicating the Scaling Law of large language models. According to disclosures by the Guizhou Provincial Big Data Bureau in July this year, China currently has about 500,000 hours of compliant real-scenario data, while commercial deployment requires tens of millions of hours, a gap of over 99%.
The data crowdsourcing business
Data reselling can happen at any stage. Devices are sent to various locations, passing through staffing companies, ground promotion teams, and collectors. Any party in the chain may keep a copy before delivery, take out the memory card, and make an extra copy. To prevent data from being copied, Old Tai developed his own collection headband hardware and software. Data captured by collectors is uploaded directly to the cloud and not retained in the device. The device also needs encryption. Otherwise, data may be sold by suppliers before it even reaches the client. But the collection side always finds new ways. Old Tai discovered that some people place another GoPro next to the device the client requires them to use. Data captured by the required device is uploaded directly to the client, while the video in the GoPro is kept by the collection team. Another approach is to have collectors wear devices from different companies at the same time. A headband from one company on the head, and a phone and multi-lens camera from another company on the body. One person performs an action once and collects for several companies at the same time. Later they found that the headband was too heavy and different cameras might capture each other, so this approach was not continued. Peers came up with a simpler method: tie two phones together. The original order is completed first, and the other phone records an extra copy along the way. "Just tie them together with a rubber band," Old Tai said.
In this industry chain, the body manufacturer may not be the final buyer either. A body manufacturer made it clear to Old Tai that the company was preparing for a listing, and the data it purchased would be resold, which could become one of the company's revenue sources. Data vendors also buy data from peers, reassemble it, and sell it to the next company. The boundaries between clients, suppliers, and intermediaries in the industry have become blurred. To avoid data duplication, client body manufacturers have also begun "database collision" checks before procurement, using algorithms to check whether data duplicates existing data. Most transactions also have no advance payment. Data vendors deliver samples or full batches first, and body manufacturers decide whether to pay after spot checks. The price of monocular data has already been pushed down. Two months ago, it could still be quoted at 50 yuan per hour; now the quote has dropped to 30 yuan, with actual transaction prices at only 15 to 20 yuan. Data of unknown origin is priced as low as single digits. Some data companies are shifting toward collecting binocular data for clients. By the end of the year, various companies may each have thousands or tens of thousands of hours on hand. As long as someone starts selling, others will follow. Low prices do not necessarily mean the data can be used. Some take advantage of information gaps to sell overseas open-source data with highly blurred video. Irregular collector movements or incorrect camera positions can also render data useless. This is a business of organizing scattered labor to produce data and earning the price difference through layer after layer of subcontracting. It scales quickly, but prices, orders, and data flows are all difficult to grasp.
Data vendors are becoming increasingly marginalized
Old Tai calls robot data a "thankless business." The company has fewer than a hundred people but must manage collectors, devices, and projects scattered across the country. After an order is signed, it still has to handle personnel training, equipment transportation, action design, data upload, labeling, quality inspection, and rework. Many of his collectors come from small shops in county towns, and the shops are natural collection scenarios. Ground promoters ask shopkeepers every day: "Want to make some money?" Many people think they have met a scammer. If the pay is too low, no one is willing to do it. A shop with thousands of yuan in daily revenue means Old Tai must settle at least at 30 to 35 yuan per hour. After collecting 100 to 150 hours, he gives the shop owner another three to four thousand yuan. To avoid scenario duplication, that shop cannot be used again, and the ground promoter must find the next one. A task subcontracted downward at 50 yuan per hour passes through two or three layers, and collectors ultimately receive twenty or thirty yuan, with each layer taking 5 or 10 yuan. The price of ordinary monocular raw data has fallen to 15 to 20 yuan per hour, while data with gloves, binocular footage, and labeling can still sell for 150 to 200 yuan, but clients complain it is too expensive. Model companies are taking back cleaning, labeling, and quality inspection work. Model companies have their own people and computing power, and they also know better what format of data they need. They have begun keeping cleaning, labeling, and processing in-house, handing only collection tasks to external companies. Data companies are left with only finding people, distributing devices, and managing delivery. Companies organize more and more collectors but may not be able to retain data and clients.
More than current prices, Old Tai worries that demand may suddenly stop. He judges that after concentrated procurement by body manufacturers this year, by the end of the year they may have accumulated hundreds of thousands or even millions of hours of data. In the first half of next year, these companies will need three to six months to clean data and train models. During this period, new procurement may pause. If the models still show no progress after data increases, the rationale for procurement will also disappear. Data vendors and body manufacturers rise and fall together, and he is very worried that clients will run out of money. Most buyers are robotics companies preparing for listing. The robotics industry is now undergoing intense elimination. If these companies fail to list, return to primary-market fundraising at a lower valuation, or directly cut budgets, data procurement will be directly affected. At a cooperation gathering, Old Tai and several data vendors talked from taking orders and making money to whether this business could survive until next year. Several people worried that models would not train well, body manufacturers would not list, and data vendors would not make money. They still hope everything will improve, that data scale will continue to expand and capacity will continue to grow.