The Endgame Of Vertical Integration
Pareto Frontier: Model-Harness Co-Design
Model Labs and Agent Labs are converging.
Anthropic continues to develop first-party applications that compete with their customers..
Whilst Harvey comes full-circle and starts training models..
Workload-Harness Fit provides a framework to assess which workloads are best optimised through either pure agent engineering (e.g. a harness that doesn’t touch the weights) on the one end through to end-to-end (pre)training runs on the other end.
Legal AI workloads have the following properties: modest volumes, high value per task, medium on verifiability, medium-duration time horizon per task. In sum, legal workloads would sit somewhere in the middle on this spectrum: the case for model training isn’t clear cut.
But the training imperative has never been clearer.
The enterprise AI market has stabilised and settled on intelligence per $ as the cleanest ROI metric that resolves disparities in headline numbers like input/output token costs, token-intensity, and more broadly benchmark maxing on tasks that don’t reflect the ‘untrainable’ work inside enterprises.
A Pareto Frontier has been established, with buyers assessing vendors’ ability to continue pushing it further out through optimising model-harness combinations.
What is a harness, anyway? There’s a running joke that the pejorative ‘wrapper’ term was just rebranded to ‘harness’, but the reality is it’s simply an evolution of what infrastructure needed to be built to realise the full potential of raw intelligence.
This anatomy from Langchain helps:
Models (mostly) take in data like text, images, audio, video and they output text. That’s it. Out of the box they cannot:
Maintain durable state across interactions
Execute code
Access realtime knowledge
Setup environments and install packages to complete work
These are all harness level features. The structure of LLMs requires some sort of machinery that wraps them to do useful work
That was the premise of agent labs: build the best harness around base models for a given domain.
That co-existence is under threat.
Anthropic is co-designing the harness for their models.
I think I am quite biased but I also think that it is impossible to get the maximum possible performance without tying together the harness and the model.
Now the components of the harness and maybe the thickness of the harness will change over time as models get more and more capable.
However, when we test our models and when we are assessing their performance we always have to test it in conjunction with a harness.
And are we going to test it with all the different harnesses of the world? We're going to select the harnesses that we have built.
And so there is an aspect of the necessity of building models is that you have to be testing them with harnesses, and that sort of keeps them paired together.
Yang Zhilin, co-founder of Moonshot AI, the lab behind the Kimi family of models, sees this as the next natural evolution of the company’s push into first-party apps:
What it essentially does is reverse-engineer the model’s training process. Because the model’s training process also relies on all kinds of means — you can imagine that Anthropic trained such a model using its in-house environments, tools, and scaffolding, but it didn’t open those directly to you.
Through reverse engineering, you get closer to fitting its distribution — which tools work best? Which system prompt works best? What kind of context engineering works best? It is a process of reverse engineering.
But you’ll find that when a model company builds a “first-party product,” the logic is completely different.
You no longer need that reverse-engineering process; it becomes a forward approach. I design the tools first, design my context-engineering methods first, and then I train the model inside this very environment — so the model naturally performs better in your environment.
These are two different lines of thinking, but the second one probably has a higher ceiling.
You can integrate tools and models much better. If the model handles something poorly, you can adjust the tool design, make it better, and at the same time train end to end. This is also a fairly large variable in the way development is done.
Poolside’s Eiso Kant has focused its harness development on coding and long-horizon software engineering tasks, eschewing tool registries and MCPs and effectively reframing the definition of a harness to be container, filesystem, provisioned credentials, and codebases:
You already see this happening more in models because when you start training them in RL, the models wanna be free. They wanna be able to do the thing they wanna do in the most efficient possible way, and it is not calling one of the 50 tools in their like system prompt.
And so I’m a very big fan of give the model a minimal harness, as minimal as possible, give it a container in which it has its own code base, right? The, got a models code base that has access to the API keys and data sources and little libraries and documentation that it needs, and just let it run free at the task. and I think that is the way we’re going. I think we will, in 12 months, not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.
Whilst conceding there is room for companies to build bespoke harnesses for specific capabilities:
Foundation model companies with their harnesses will really push them because it’s just operationally, the best way to have scientific rigor in improving your models.
But also someone who takes our model and really does a lot of work on improving a harness is going to compete us, as they should. and that’s just because the harness is the stopgap between what the model is capable of and what it needs as additional instructions, and what it needs is access to data and tools, right? And that’s ultimately, I think, what a harness is.
As you build more capable models, you’re improving the instruction following the models. And so additional harness is just saying, “Hey, if you encounter X, Y, or Z, behave this way.” And so even if you would say that two models with two different harnesses can equally reach the same capability that you care about, a harness that is really tailored towards a capability will do it more efficiently.
Until now, most model labs had only built harnesses for specific domains, primarily coding. The intent to expand surface area is clear from feature releases and M&A activity.
The gains from co-designing models with harnesses is why agent labs need to move into model training, or risk obsolescence when the model labs concentrate their resources on your domain.
As I’ve said before, the economics of training have changed dramatically over the last four years, which means that becoming a ‘lab’ no longer has the same connotations of capital-intensive, open-ended R&D that it once did.
Pushing out the Pareto Frontier for a given domain is the most important mission for agent labs to work on.
Our goal is to offer, at every point along the Pareto optimal frontier, the best option for intelligence and price.
That might end up being the most durable differentiator versus model labs.
Anthropic’s inference gross margins on its API business are reported at over 70% currently, up from roughly 38–40% in 2025.
The challenge is for agent labs to offer the best intelligence per $ at margins that the model labs can’t match.
When the model labs are sitting on lucrative API or advertising businesses and a hundred competing priorities, their ability to concentrate the resources needed to match those margins is slim. Agent Labs should be the ultimate winners in many vertical markets, for these reasons and others (e.g. regulation).
To offer the Pareto optimal frontier for legal, finance, or other domains, agent labs need to co-design models and harnesses too, which is what’s now unfolding.





