Writer's AI Harness Cuts Token Costs 52%

The enterprise has an AI cost crisis, and it's not the model's fault
Enterprise AI spending is going in one direction, and it's not down. Frontier labs charge $5-$15 per million output tokens for their most capable models, and as agents grow more autonomous, handling multi-step workflows, calling tools, and iterating on results, each task consumes more tokens than ever.
But the problem isn't just the model's price tag. Writer, a company that builds AI tools for marketers and revenue teams, argues the real cost driver is something most customers never see: the harness.
"Everyone's been chasing the next benchmark," CEO May Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that."
Writer's answer landed Thursday: Palmyra X6, a new flagship model post-trained on Z.ai's open-source GLM-5.2, paired with a significantly upgraded agentic harness. The combined system cuts costs by 52% on average, with a 48% improvement in speed and a 10% quality improvement, according to the company.

The harness is the hidden lever
The harness is the infrastructure layer that mediates between an AI model and the tools, data, and external systems it needs to do its job. It decides which tools to call, how to present results, and how to chain subtasks together. Writer's research found that small changes in harness efficiency were a more reliable way to reduce costs than model choice, cutting costs by an average of 40% across testing.
"The harness is the one component whose efficiency multiplies across every model an organization runs, present and future," the researchers wrote. A model swap saves a fixed percentage on per-token price. A harness optimization saves that same percentage on every token, every task, every agent, forever.
The upgraded harness launches with Playbooks and Skills, repeatable agent workflows that let teams build and share standardized processes. A dedicated analytics dashboard gives administrators visibility into consumption, performance, and spend across teams, with configurable alert limits.
The numbers, and how they stack up
Palmyra X6 itself is competitive on price-performance. On nine evaluations, it scored an average of 0.87 at $2 per million input tokens and $8 per million output tokens. That beats Claude Opus 4.8 (0.86 at $3/$15), GPT-5.5 (0.80 at $5/$15), and Gemini 3.1 (0.77 at $2.50/$10), according to Writer's benchmarks.
The model completes tasks in 26 seconds on average, produces 82 tokens per second, and can operate autonomously for up to eight hours without drifting off task. Writer also claims it is the least politically biased model the company evaluated, a consideration for marketing teams that need to reach diverse audiences.
Palmyra X6 is model-agnostic: it sits alongside other models or imported models from Azure and Amazon Bedrock, giving customers the flexibility to route tasks to the cheapest or most capable model for each job.
A broader distrust of AI labs
The cost crunch has a second dimension, and it's cultural. Enterprise CIOs are increasingly skeptical of the major AI labs' incentives. When a lab that charges per token also controls the model architecture, the hardware, and the pricing, the customer has no leverage, and no guarantee that the lab's interests align with keeping costs down.
"The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," Habib said. "The labs don't deeply understand how to help an enterprise get benefit from AI."
The same week Writer launched Palmyra X6, DeepSeek raised some of its V4 API prices by more than 10x, citing capacity strain. The coincidence is a convenient illustration of Habib's point: when demand surges, the lab that controls the pricing has the power to pass on every cent of the strain.
The bet that the boring stuff matters
Writer's Palmyra X6 strategy is a bet that the next phase of enterprise AI adoption won't be decided by which model scores highest on a benchmark, but by which vendor can make the economics work at scale. The harness, the invisible, unglamorous infrastructure that routes tokens, manages tool calls, and surfaces costs, is where that battle will be fought.
Co-founder and CTO Waseem AlShikh put it plainly: "The enterprise wants token consumption to explode, it means adoption is happening, but they need costs to flatten."
Whether Palmyra X6 delivers on its claims will depend on real-world deployments, not benchmarks. But the thesis, that the harness matters more than the model, and that the labs' incentives are at odds with the customers', is a genuinely useful lens for understanding where the enterprise AI market is headed.
Sources
- Writer introduces new AI model and upgraded harness to contain token costs, TechCrunch
- Writer launches major agentic AI improvements with Palmyra X6 flagship model, SiliconANGLE
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity, InfoWorld
- Writer says its new Palmyra X6 model cuts AI agent costs by 52%, VentureBeat