Most people assume the hardest part of building a frontier model is the model. The architecture, the training run, the research. After a year sitting on the other side of the table from those teams, I have come to believe the opposite. The architecture is rarely the bottleneck. The data is.
The gap between what data vendors think labs want and what labs will actually pay for is the most expensive misunderstanding in this corner of the industry. Here is what I have learned from closing those partnerships.
The constraint is data, not design
The research caught up to this a while ago. DeepMind’s Chinchilla work showed that most large models were badly undertrained for their size, and that you should scale data and parameters together rather than just making the model bigger. Around the same time, Epoch AI’s analysis of whether we will run out of training data projected that the supply of high-quality public human text could be effectively used up within this decade, pushing the field toward synthetic data, other modalities, and squeezing more out of every example.
Translate that out of paper language and you get a simple business fact. The frontier is moving from “get more data” to “get the right data, structured correctly, delivered reliably.” That shift is where a good data partner either earns its place or gets quietly dropped.
And yet most vendors are still selling volume.
What labs actually ask for
When I sit down with a lab’s research and infrastructure teams, almost nobody opens with a request for more raw data. The pain points they describe are operational, and they are remarkably consistent across teams.
The first is annotation consistency. A multimodal dataset where labels drift in meaning from one batch to the next teaches the model noise. Two annotators interpreting the same instruction differently is not a rounding error at frontier scale; it is a measurable drag on what the model learns. Labs want a partner who can hold a labeling standard steady across hundreds of thousands of examples.
The second is delivery cadence. A research team plans a training run against a calendar. If your data shows up late, or in unpredictable bursts, you have not given them a dataset, you have given them a scheduling problem. Predictable, ahead-of-schedule delivery is worth more to them than a larger one-time dump, because it lets them plan the thing that actually costs money, which is compute.
The third is format interoperability. Data that does not slot cleanly into an existing training pipeline creates engineering overhead the lab was specifically trying to avoid by working with a partner. If they have to write a week of glue code to ingest what you sent, you have added cost, not removed it.
None of these are glamorous. None of them are what a sales deck leads with. All of them determine whether the dataset meaningfully improves the model’s reasoning and comprehension, or just inflates a row count.
Why most vendors miss it
The miss is structural. Selling volume is easy to pitch, easy to price, and easy to demo. Selling operational reliability is hard to put on a slide, because it only shows up over weeks of actual delivery.
So vendors optimize for the sale instead of the training run. They lead with the size of their data pool and the polish of their deck. The lab, meanwhile, is trying to answer a quieter question: if we wire this partner into our pipeline, will the next training cycle go smoother or rougher? A big number does not answer that. A track record does.
There is a second structural miss, and it is about how trust gets built in this market.
You earn a frontier lab before you pitch it
Frontier labs vet vendors rigorously, and informal reputation carries disproportionate weight. A warm word from one research lead to another moves more than any amount of outbound. In that environment, a relationship-first approach is not a soft nicety. It is the only thing that works.
The partnerships I have closed did not start with a proposal. They started with a peer-level technical conversation. In one case, the channel opened because a key stakeholder and I shared an academic background, and that single thread of common ground turned a cold vendor conversation into a candid one about where their pipeline was actually hurting.
What followed was capability shown before anything was asked for. We hosted an on-site, full-day technical review, walked their team through our data-engineering infrastructure and our delivery benchmarks, and let the work speak before any commercial terms came up. The first dataset went out inside three weeks of first contact, ahead of schedule. That delivery, not the deck, is what earned the internal advocacy that carried the rest of the relationship.
The sequence matters: credibility, then candor about the real problem, then demonstrated capability, and only then a commercial proposal. Reverse it, lead with the proposal, and you signal that you are optimizing for the close rather than for their model. Sophisticated teams read that instantly.
What I would tell anyone selling into this market
Lead with the operational truth, not the volume. Annotation consistency, delivery cadence, and format fit are your real product. Say so.
Demonstrate before you propose. An on-site review or a live walkthrough of your pipeline does more than a quarter of follow-up emails.
Treat delivery as the pitch. The first dataset you ship, on time and clean, is worth more than anything you said to win the deal.
Build the relationship before you need it. In a market this small and this reputation-driven, the warm introduction is the whole game, and you cannot manufacture it the week you need to hit a number.
The labs building the most capable models in the world are not short on ambition or talent. What they are short on is partners who understand that, at the frontier, the unglamorous operational details are the product. The model is the part everyone watches. The data pipeline is the part that decides whether the model gets better.
n