The Human Data Economy
The industry building AI’s human data pipeline is booming, human error and all.
Amy Smith
Rebecca Milde

Job postings data offers a timely view into emerging roles and areas of the labor market that are difficult to capture through traditional workforce data. They can also reveal not just where demand is growing, but how employers are defining new areas of work through skills, responsibilities, and qualifications they seek.
The following article, by Amy Smith, uses Lightcast job postings data to explore the emerging market for “human data”. By tracking postings referencing this phrase, the analysis surfaces how frontier labs are building this part of the artificial intelligence ecosystem. The growth of these postings offers one window into the rapidly evolving AI labor market. Looking specifically at job titles, demand for AI Trainers, like AJ in Smith’s article, is also expanding. In the U.S., more than 250 companies have advertised AI Trainer roles over the past three years. These postings move quickly, remaining online for a median of just 11 days, well below the 24 day average.
As AI development continues to evolve, job postings can provide an early signal of how the work behind it is changing. The article below takes a closer look at what those signals reveal about the growing human data economy.
– Rebecca Milde, Lightcast
******
On a quiet Thursday morning, I logged into a zoom call with AJ, an AI trainer who works with human data companies like Outlier and Turing. AJ is in his final year of a Data Science Masters program in the UK. “AI trainer” is a relatively new job title that refers primarily to independent contractors producing training data for AI models, evaluating model results, and feeding them back into model training. AJ’s AI trainer gig is a part-time income supplement. He says that contracts are typically 1-2 months. When I asked if he knows which models his work contributes to, he said no, it’s confidential, but it slips from time to time.
Meanwhile in San Francisco, a recent Anthropic job posting is looking for someone to sit at a strange new intersection. "Anthropic's Human Data Platform team builds systems designed to collect data that improves our models. This includes the infrastructure to emulate real-world environments and tasks, novel interfaces for data vendors to use, and the pipelines that enable researchers to gather high-quality data at scale. As Claude's real-world usage evolves, so do our data needs — and our tooling has to keep pace." Anthropic is looking for someone to manage the human data pipeline that AJ’s work is feeding into.
Roles like this, that focus on managing human data to teach AI models to do what humans do, are on the rise. "Human Data" jobs have grown from a handful of postings at the beginning of 2022 to nearly 300 cumulative postings by July 2026, according to workforce data from Lightcast, which tracks job listings over time. Every major frontier lab is hiring for people who can design research questions, define human data needs, source domain experts, build data pipelines, and turn what those experts know into something a model can learn from.

Why the focus on human data when the AI industry's own leading thinkers have argued we're leaving the age of human data behind? Richard Sutton and David Silver, two of the field's most prominent researchers, call this period the "Era of Experience" (the title of a preprint chapter forthcoming in Designing an Intelligence from MIT Press). The idea is that in this era, the next leaps are anticipated to come from models learning via their own trial and error, rather than by imitating human language. Pre-training data, after all, is a snapshot in time. If AI systems will eventually be engineered to learn continuously on their own, human data begins to look like a crutch for frontier labs that still need to bring in revenue through software that imitates what humans do faster and better than has been possible in the past.
But the job postings don’t tell the whole story. In practice, labs aren't choosing between human data and machine experience; they're building a hybrid. LLM-based agents learn within simulated environments that humans design, or in the case of Turing’s Project Lazarus, from startups paid up to $1M for their code, communications, and business records. Sometimes the data comes from defunct companies, as was the case in Google’s purchase of Spirit Airlines data for $10M. The promise being that AI agents can be given access to these environments to learn through doing in a “real world” setting. Maybe it’s not human data vs. experience afterall, but somewhere in between.
The Tacit Dimension
In 1966, the philosopher Michael Polanyi published The Tacit Dimension, arguing that "we can know more than we can tell." According to Polanyi, some human knowledge is developed only through experience. You can read every book on piano technique ever published, but good luck sitting down and trying to play a Chopin étude on the first try. A gymnast who’s spent years perfecting a back handspring might be able to describe to you how to do it in words, but it’s even more effective for her to show you and say, “you do it like this,” then watch you do it and correct. Little by little, you try, fail, receive her feedback, and learn.
The first wave of large language models was bootstrapped by the knowledge that humans could express. What’s missing is the layer Polanyi described: the tacit dimension. It’s how a sound engineer can translate a request to make a sound more “floaty” into something his producer is happy with. It’s how a veteran transportation planner has a gut feeling for how the public will respond to a proposed bike lane. This tacit dimension is intuitive, and it’s where frontier labs are now focusing their attention.
Yet teaching a model how to be intuitive is difficult. AJ told me that the human data companies he works with have been providing increasingly detailed rubrics. The specificity provides standardization and consistency but makes the whole process rather prescriptive. It doesn’t leave much room for intuition. When I asked AJ if he felt like the end goal was to train models to eventually be capable and intuitive enough to replace him, he laughed. “They can’t do that. It’s impossible. They need us.”
The job postings and massive recruiting efforts for AI trainers suggest AJ’s right. Even if models are spelunking in real-world data and virtual environments to learn new skills, someone still needs to craft a testable, measurable research question, structure the inputs and outputs, and analyze the results. Legal experts need to ensure privacy and IP protections. Security engineers need to control for real cybersecurity risks.
The Market for Human Data
Human data job postings from the frontier labs and tech companies call for hybrid skill sets with general AI literacy and a working grasp of generative and agentic systems. Some postings call out familiarity with alignment techniques and feedback channels, showing a priority for ensuring models are aligned with human preferences and goals. A competitive candidate will have experience managing pools of domain experts (data scientists, lawyers, doctors, investors, marketers) and turning their judgment into training signal.
How big are these pools of experts, exactly? How many AJs are there out there? The data on the exact number of humans involved in the creation of human data is limited. If funding is a proxy, 2025 was a record year for money raised by human data sourcing companies. More than $800M was raised collectively, which doesn’t include the 49% stake in Scale AI acquired by Meta who invested $14B in the training data startup.

It’s worth noting that companies like Turing and Handshake didn’t start offas human data sourcing companies, but rather talent recruiting companies. Turing made the switch after a meeting with OpenAI in 2022 led to the realization that the company was sitting on data gold. Handshake made the move to AI training data in 2025. The pivots are paying off, and new entrants are popping up. Deedy Das, partner at venture capital firm Menlo Ventures, estimates a ~$8.5B revenue run rate and ~$100B in valuation for AI training data and RL environment companies as of July 2026.

An Organized Mess
Interacting with an LLM can feel mysteriously on-point. Training one can seem equally as uncanny. Create all the rubrics in the world, sometimes model training and fine-tuning results remind us the way humans learn isn’t so linear. Or as Sutton points out in a September 2025 interview on Dwarkesh Patel's podcast, “Supervised learning doesn’t happen in animals.”
In August 2026, researchers at Surge AI trained a model on simple office work, and it got better at coding, even though there was no code in the training data. The researchers compared it to a junior engineer who takes a year off to plan weddings and comes back as a better code reviewer. Not because they were doing any coding in their time off, but because they learned how to lead a complex project with lots of dependencies.
This past May, the team at Anthropic wrote that they wanted to improve AI’s ability to resist acting out of alignment. Training the model on examples of doing things “right” had its limitations. However, when they taught AI ethical reasoning and fed it stories portraying an AI aligned to Claude’s constitution, misalignment was reduced by more than a factor of three. It didn’t matter if the evaluation scenario at hand was related to the stories. The model was able to generalize its learnings once it understood the “why”.
Reports like these give the impression that training AI models involves a good amount of guesswork, and it’s important to not be too attached to the original goals, which are seemingly endless. With so much hypothetical spaghetti to throw at the AI training wall, having human data nearly at the push of a button can feel like a competitive edge. On the flip side, without any constraints, labs run the risk of pushing ahead on too many fronts, unable to take the time to carefully assess the quality of a dataset or interpret the implications of the results. The human data manager is a cat wrangler, aligning teams on research questions, pressure testing methodologies, and ensuring the right data is coming through the pipelines. They’re responsible for creating order from the chaos.
Still Human
AI model development is, for now, a human-intensive data operation: recruiting experts, watching them work, gleaning their knowledge, and grading machines against their judgment. That said, and humans being humans, it’s not immune to corner cutting. When AJ mentioned that his work is evaluated and scored at the end of a project, I asked who did the scoring. He said it was usually anonymous. I asked if he thought it was possible that an LLM graded the outputs. He smiled and said that he suspected so, describing the system as “cannibalistic” at times.
This summer, I personally began receiving emails from AI training data company Mercor telling me about opportunities to earn money as a data science expert. Curious, I created an account. The interview was conducted via an AI agent’s synthesized voice on a video chat. The questions flowed organically toward whatever came up in my response to the previous question, even if it didn’t match the subject area for the role at hand. There was no indication of whether a human or another AI system would determine my fit. It makes sense that companies specializing in the creation and curation of training data for AI models need to use AI to be efficient themselves. But it begs the question: who’s checking the AI checking the humans? At the end of the day, there’s a huge amount of human data being generated, and somewhere between humans asking for it and the models learning from it, there’s a lot of room for human (and machine) error.
As our conversation came to an end, I asked AJ about his thoughts on the experience vs. human data debate. “Sometimes it feels like they’re just copying our responses,” AJ told me. “Making it seem like the chatbot is making something new, but it’s regurgitating. AI either becomes a chimera of 10-15 workers or a person who learns something.” While no one outside of the labs has the visibility to know which one we’re closer to, both visions are fueling the new economy for human data.
by Amy Smith
View other work of hers here
Related Posts

Q: What is a Forward Deployed Engineer? A: The fastest growing AI job.

Degree Requirements are Dropping—But They’re Still Higher for AI Jobs
