Intent and a new class of PMs
Upskilling Product Managers for new data paradigms
By Sean Cai
More Private Data Markets Comps/Analysis on Substack or on seancai.com
This is a short companion piece to my State of Data on Jan 2026 which is undergoing some final revisions before being published tomorrow.
In this piece, I explore how to systematically identify what makes an AI app layer market mature, a data-centric perspective on the interplay between human data and RLaaS companies in getting app layer markets there, and a new JD/skillset that will probably evolve into a new job title in AI product companies.
Preamble
Human data companies sell to labs because labs are ML talent rich, but data poor. RLaaS companies sell to enterprises becuase enterprises are data-rich, but ML talent poor. That’s not to say that labs can’t build out their own human data teams (they certainly do) or that human data companies don’t have ML expertise (they often need to train on their own data to prove its worth) but these core competencies explain why our AI app layer products are the way they are.
There are few companies that do both extremely well. Cursor and Claude Code are products whose companies both are large vectors of data via usage as well as being able to hire large amounts of MLE talent. In such a way, you can reach escape velocity via products - you no longer necessarily need human data companies and external RLaaS expertise at scale because your core business generates PMF and enough usage to attract and hire talent.
We generally also recognize this as the definition of PMF in AI app layer products. Even if you continue to have negative margins, you are well funded enough to invest enough in R&D/product advantages to beat out the competition.
If we think this way, then the interplay of RLaaS and human data companies serve an ironic unique role. One imagines that, if they don’t seek to satisfy their equity investors’ expectations of empire building, then they will instead opt to enter the model/app layer themselves. This is an uncomfortable reality for many RLaaS companies, who instead opt to delay the inevitable by oscillating to more RLaaS use cases (sometimes vertical specific) in a bid to build more horizontal infra.
In reality, the shift to an app layer/model company here is gated by limitations regarding talent aggregation (and by extension, getting to be funded like a neolab). Many RL env companies are even forthright about this in initial investor conversations - we must sell data to labs or do RLaaS engagements today in order to get the prerequiste funding/trust with researchers/runway to eventually do this ourselves. It is the ultimate and purest form of learning on the job.
You may look on this with some sort of fatalistic skepticism, but I’ll remind that neolabs are not the only modality of success here. A quarter of them that execute well will probably be folded/bought out by a variety of empire-building general purpose data cos, generalized RLaaS builders, and even AI app layer cos who’ve attained PMF that want more talent. Another quarter will be able to successfully make the leap to domain specific app layer winners.
This is an important point for my friends unfamiliar with data markets who are scratching their heads at why investors are funding scale-type business models with little longevity. None of these companies want to stay this way - this is simply the best plan in the current day to build a talent-aggregating, revenue-generating service product to learn how to build best in class AI products and infra.
But this is not the purpose of this article; I want to describe the difficulties with adapting to RL and automation adoption for non-coding related white collar work domains.
Mature App Layer Markets
The coding use case for the AI app layer had so many natural advantages to be the first mature AI app layer market, that one could compare its blessings to that of the geographic blessings of the US. We had coding data in naturally model actionable and readable formats amicable for reasoning trace analysis in the form of github repos. Coding is easily verifiable and data available in modern mass (with rewards fairly objective). Coding data is abundant in every company today, and unreasonably available online with something called “open source” which puts an inordinate amount of enterprise data online. Coders and consumers of a potential new age product are willing early adopters by nature of their job and incentives for app layer solutions require the least amount of engineering work between a base model and its actual application (only a VSCode fork with a prompting sidebar, as some in early 2023 would have said).
Its no surprise that coding was the first mature AI-adoption market, given how amicable these conditions were to any sort of AI-paradigm that was RL based (based on RL’s propensity to be a data hungry algorithm benefitting from clear rewards).
These conditions do not exist for other white collar work domains and use cases. The closest equivalent today is search, but even search lacks the “open source” abundance of coding and has its early adopter users spread out across multiple verticals. Search data is, however, abundant, with reasoning traces clear (if tracked) as well as verification most aligned with how models already do next token prediction.
If you want to be a systematic and extrapolative thinker, you can develop a framework around which AI app layer markets are next to mature. Those that share as much similarity with the conditions described with coding are probably ripest for maturation. The most important condition being data availability in some sort of standardized, reasoning trace heavy format, in absence of some sort of ML paradigm breakthrough such that we aren’t so RL and reward function dependent. On that last point - I can only imagine that breakthrough being long horizon memory related.
Our checklist for model companies and app layer solutions that one can use to determine maturity timelines and opportunity size is:
*The unit level refers to individual tasks, not entire workflows. This is important as task correctness varies wildly in some white collar enterprises due to shifting, un-ingestible best practices and tool calls (while coding envs may, say, have different packages coming in and out, they are generally easily digestible or already included in pre-training context by models).
Theoretically then, in absence of an unreasonably large incentive for a large enterprise domain specific company in that area to commit large resources to build a domain-specific neolab, human data and RLaaS companies can use this framework as a guide to figure out which industries represent the largest TAMs to pursue.
RLaaS today:
The views of most human data companies and RLaaS companies are mostly actually not that divergent or unique.
At the highest abstract level, they all revolve around working with customers with the most sophistication, with the highest budgets for experimentation, with the most high fidelity data lakes, with the most sophisticated workflows, to learn the most to find recurring use cases.
As you can imagine, they all come into this endeavor with bits and pieces of hypotheses and alpha from prior jobs, but nobody with a cohesive platform idea. This is a good thing as there is, resultantly, still enormous white space. This also makes them incredibly market beta levered (as infrastructure investments often are).
There are two classes of RLaaS customers today:
Sophisticated cos who throw hardest product MLE problems
Doordash, airbnb and Clay are such companies. Products that borderline on needing skillsets of their own (like Clay) have huge product surface area for robust agents that automate workflows, reduce user onboarding friction, and who can benefit greatly from post-training. Gig-adjacent models like Doordash and AirBnB have incredibly rich RL envs with lots of high fidelity interaction data (see Doordash restaurant management use cases) where RLHF is apt for product improvement on very deterministic KPIs (eg 5 star ratings).
But these contracts are also heavily customer favored. I often see RLaaS companies contract with these companies, whose names are the most fought over as of Jan 2026 by RLaaS cos, with negligible service-based fees for the initial discovery period and no promise of procurement whatsoever. Solving the RL problems of the likes of a Doordash or Coinbase is something to brag about and launch productized momentum (see Forge) but also the most competitive customers to win currently in this space.
New AI app layer companies desperately trying to breach escape velocity
Harvey, Rogo, and equivalents with widely used app layer. Once escape velocity pmf is hit, they can theoretically go the way of Cursor, and generate enough user data for product improvement, as well as hire the MLE talent needed to solve their hardest model problems.
There isn’t a strong argument or conclusion I have as to whether to serve either or focus on one. The 3 month pilots that both of these types of companies employ are still ongoing, and their KPI metrics for conversion still dubious. Many of those pilots also don’t have clear ARR expectations for conversion (is there even a way to make this recurring revenue) and also are set up with the expectation that the customer would simply seek to acquire the RLaaS company if true value were to be had.
The RLaaS name is actually a terrible misnomer, probably coined during a time where we thought RL was as much a paradigm shift such as to merit a distinct category from MLE infra. As mentioned before, RLaaS companies currently do not spend most of their time doing RL, but rather setting up organizations’ data ontologies to be amenable for sophisticated post training.
A new class of PMs (model managers?)
Two classes of problems exist with creating net new datasets for the increasingly unverifiable white collar domains we care about:
Intent is hard
Discerning user intent and converting this into actionable reward signals for model training is exceedinly difficult.
Old PMs talked to users and then converted them into product data for product development. Entire interview guides and standardized roles evolved around this (APM programs). The natural next step here should be PMs who design products in ways that are amenable for model-actionable data collection from users.
The right reward rubric for a vague white collar workflow in absence of robust pre-existing reasoning traces in web 2.0 software applications must be discerned via a combo of semantic user feedback and user observation. Some startups in SF choose to treat this as a product problem, but I believe the solution for this role is actually a new JD/role, which will in 10 years, become the new APM programs of big tech. The intent is too scattered and ungeneralizable, the problems too human facing across different modalities, and the context too culture based for any sort of non-human tooling to solve.
Intent changes, and our reward functions must change alongside them
The white collar domains we increasingly deal with today don’t have GitHubs, standardized industry practices (to the extent that coding does), and are subject to rapid C-suite directive changes. What this means at the MLE layer is dynamic reward rubrics. While we could focus on building the most coveted of Antikythera mechanisms here (converting messy unstructured enterprise signal into RL parameters and rubrics quickly and accurately), the more likely solution is to simply hire someone who can translate directives into deployed models quickly.
Today, this is a full stack eng, MLE, and product person all in one, with whom the C-suite peppers with constant strategy shifts. One might call that a “CTO,” but the job functions of the MLE gets ever stretched with new model paradigms and training practices. The persons with that skillset, in absence of significant equity compensation, will only become harder to find if we do not birfurcate these JDs (or at least upgrade our PMs to these practices).
The envs and tools change frequently
White collar domains in finance, health, law, (and much more so for all the domains after that) have workers with an incredibly fragmented stack. CRMs, outbound sales tools, data lakes, data sources, and productivity tooling face procurement tasks every day. Teaching models to generalize across this tooling sometimes doesn’t even work; computer use agents still struggle to reliably navigate the interfaces that different productivity tooling has.
You could build and RL env and eval set for one workflow, only to have your client change a subset of tools and break what you’ve built easily. An extremely SF-focused product solution for this subset of problems probably centers around updating RL hyperparameters and reward rubrics with agents aware of procurement-level decision making in users. More likely, this job task should lie in the JD of a new class of model-aware PMs.
A new JD
The JD that best addresses this class of problems is also one that is customer facing, and thereby has good data taste. Its a JD already loosely expressed in orgs like Decagon as “Agent Product Managers.” But the best and earliest candidates for this JD will have developed certain skillsets that make them excel:
As many of the service-based enterprises they work with will have scarce data representations for the problems they want to solve, they will ideally be well versed with good synthetic data generation and generally how to make hard and diversifiable synth data for model training, also knowing how much representative data you also need
They will recognize how much completely generated synthetic data can hillclimb a custom model deployment, likely by already recognizing how well a base foundation model does at relevant enterprise tasks
They will recognize the fallacies of what benchmarks fail to measure and how to reconcile general purpose public benchmarks against internal product evals
They know how to acquire data from real world stakeholders without saying “I want to buy your data” in creative solutions that show an impeccable understanding of incentives
They know how UI/UXs and product experiences can be designed to balance the interplay between “good user experience” and “captures enough reward signal for useful model training”
There is no excuse to be “non-technical” any more in the traditional sense. All products nowadays incorporate deployed models. All models that are good in production today involve some sort of post training. All post training for good product development requires a good understanding of data, how to build good UI/UXs for users that allow PMs to capture reward signals in model actionable formats, and subsequently bridge MLE and product. Tactically, this might look something like: how do I design an app layer product well so as to set up real time data pipelines to update my MLE teams’ RL envs’ reward models with new product directives?
When is synthetic data useful?
For some more non-technical readers, synthetic data at its most banal form can be prompting chatgpt to generate RLHF pairs and using that to finetune chatgpt (don’t actually do this, you will waste a lot of money in finetuning). In applied ML, it most often means some sort of sophisticated model usage to create and augment training data for model hillclimbing.
Synthetic data is good for creating random sampling with some context augmentation if you already have omni-representative datasets baked into a trained model.
Synthetic data is inherently non-useful for end to end frontier model training because pushing frontier capabilities inherently means providing out of distribution data. Synthetic data works when:
A) You have close to near certainty that the entire range of possible generations roughly is known and represented in the model’s training set/context, and
B) You have some sort of pipeline to capture net new generations as they are generated in the real world
(see task generation for RL envs for white collar work domains where human experts are still providing reasoning traces themselves).
In many enterprise model training use cases, there are actually points of “good enough” model performance to be economically valuable, as compared to our benchmark pushing in labs. This presents a wide aperture of use cases for synthetic data:
In MoE implementations where we train many small models who act as essential reliable APIs for singular workflows, condition A is often true and condition B is easily buildable
Enterprises have large stores of data traces for very repetitive white collar work samples, and ask for more reliable automations here than brittle RPA. Some of these issues are out of distribution problems already represented in base foundation model capabilities, and synthetic data gen can be used out of the box to hillclimb. It is, however, extremely difficult to know what is truly already within foundation model capabilities (especially with obscured benchmarks like Apex and GDPEval), especially if swapping models frequently.
Getting to a crucial mass of data outcomes for extremely reasoning trace sparse workflows to prototype and build a model that gets from 0-1 on benchmarks. This overcomes a lot of sparse model training issue diagnoses and allows you to hone MLE focus on benchmark climbing rather than a plethora of possible model initialization that you couldn’t rule out because of insufficient data.
Mistakenly applying synthetic data poisons models, kills training runs, and racks up compute bills. It is often hard to diagnose as well, if poor data quality via mis-applied synthetic data practices is a core contributor to stubborn maxima searches. More unscrupulous actors, whether aware or not, take advantage of this by selling some datasets to model trainers where a small portion is AI-generated. This cuts corners (and by data sellers’ incentives, can produce higher revenue growth through more data output) and human data companies’ reputations, but can be mitigated with data source auditability.
Luckily, the knowledge for good synthetic data production and application is diffusing such that they make an integral part of the Antikythera mechanisms app layer companies, human data cos, and RLaaS cos alike build in-house in pursuit of business objectives.
If one also expects synthetic data to diffuse in popularity for the aforementioned use cases, there is a large market opportunity for auditability tooling for synthetic data to reconcile its proper application, just based off of how difficult it is to pinpoint poor synthetic data application in model training difficulties. The cultural shift for proper synthetic data usage in applied ML teams and new JDs briding MLE and product teams, though, must disseminate first before this can happen.
Everything is priced correctly
The reason why most RL env startups in 2025’s seed funding rounds near exclusively were done at 30-50M post is a beta factor. As mentioned before, if you believe that RL env sales to labs and enterprises, investing in a talented group of MLEs with lab connections, while also providing some marquee AI app layer company introductions is a bet on their ability to find an out (like the outs from my crazed graph before). Most don’t actually come into raises with unique market insights or advantages; if you truly did, you’d likely be incentivized to work at a lab or be acqui-hired by one.
Similarly, credible teams working on RLaaS are also priced via beta consideration often. Pedigreed ex-researchers will have the name to be thrown bones for service-based contracts from any sophisticated hot app layer company. Top shelf MLEs who worked on enterprise RL in mature app layer companies (Parallel web systems, cursor, windsurf, etc.) will theoretically have some knowledge on how to build the Antikythera mechanisms we seek better and faster.
You’ll notice something interesting here with regards to pricing. Though not mentioned explicitly, we are pricing based on surface area to discover the problems and build platforms that unlock AI infra’s TAM. Commercial ability is largely forsaken (because many investors believe that there are only 5-10 companies worth working with here, and all of them are within their portfolio) and AE-type product expansion from converted pilots is the priority. If the likes of AC can beat out competitors and train economically valuable models that satisfy enterprise needs, and not only build out generalizable infra as well as charge recurring fees on trained models,, they will well be worth their valuation.
Private capital lockups only embolden this because pre-seed like discovery phases are easy to drag out in times of capital abundance. Burn rate is hardly a concern in AI infra and human data cos - they are often profitable and still sitting on large war chests (Mercor) where their main thought on use of funds is acqui-hiring talent. This is different for AI app layer companies who are forced to spend vast sums on data in the form of unprofitable users (they only become actual users when they become profitable). Luckily, emboldened private capital markets only make the drive to profitability easier for AI app layer companies who are also afforded new options like investing in their own bare metal, and from which new model releases are not only free product upgrades, but also N-1 model releases are free COGS reductions.
Data acquisition, in a sense, is the only thing that these companies will burn money on, whether by unprofitable user acquisition or direct payments to human data companies.
Are we really surprised that, at the end of the day, this is a gold mine for free cash flows for startup founders? Usually SF Libertarians profess individuals should ignore the macro and focus entirely on the micro. In this case, if data center buildouts account for most of the GDP expansion in the US in the past two quarters, then its seems natural that the spend trickles down into the most AI-leverred of wrapers, data companies, and finetuning services companies serving the builders of those data centers (OAI).
Afterword
The root source of all US GDP growth is hardware investments and data. With new deeptech innovations and physical chip/data center infra investments, we get more engines for experimentation. With new pipelines and Antikythera mechanisms for collecting, distilling, and refining work traces from humans in real life, we get the oil for these engines. This is the foundation for a post-information age highly educated economy being trailblazed in the US.
There is a more elegant and versatile way to say this to the wider US labor force. SF language would be “we were put on this earth to build RL environments for labs.” Ironically, while we could always use more researchers to create more optimizations at model architectural layers (unless you believed generalization paradigm breakthroughs are around the corner), then our bottlenecks actually mostly lie with creating elegant socially engineered solutions for wrestling reasoning trace data from the masses of “anti-AI” American workers.
There will be no one size fits all arrangement for bridging the data from the real world to AI land. Individual SLAs will have to signed with 30-year family businesses over regional dishes in Lousiana’s Cajun region or town breweries in Wisconsin. AI-savvy Bay Area engineers will have to study whether “pop or coke” is the right way to talk to a Midwest call center owner to wrestle for 40-year old audio datasets for model training. The core issue of unlocking enterprise data and formatting will remain messy and services based, and where the biggest job opportunity for workers laid off from AI implementations can go.
Thinking, writing, and theorizing does nothing without action. See what I’m working on to justify and substantiate my writings here.

