Blog

Going Once, Going Twice: Notes From the Floor of the AI Model Auction

On September 25, our favorite value model got delisted. Kimi K2.6 was cheap and good at extraction, and then its host retired it. That capped a summer when Kimi K3, GPT-6 Sol and Luna, and Claude Opus 5.5 all repriced the market. Joon Lee, Head of Forward Deployed Engineering, explains how his team picks models for healthcare customers: treat it as a reverse auction with a reserve price on quality. He walks through a real enrollment agent where 4 models bid, none cleared the bar, and redesigning the job changed the outcome.

By Joon Lee

9 min read·October 1, 2026

On September 25, one of my favorite AI models got delisted.

Kimi K2.6 had been our value pick since it launched in April. It was cheap ($0.95 per million input tokens and $4.00 per million output tokens on Moonshot's API), quick, and good at extraction. I built part of a customer pipeline around it. Then Fireworks, the host we ran it on, retired it from its serverless API. The notice suggested we migrate to GLM 5.3 or Kimi K3.

Our staging environment found out on September 28, via a wall of 404 errors. (Our automated tests caught it first. I'd love to tell you a human did.)

I run Forward Deployed Engineering at ACTAVA.ai. My team works where our platform meets our customers' real healthcare workflows: care management enrollment, revenue integrity, prior authorization, and cash posting. When a model changes its price, changes its behavior, or disappears, we get the call. This summer, the phone rang a lot.

The summer the ticker wouldn't sit still

Here's 2026 from my desk:

  • April 20: Kimi K2.6 lists. Our value pick.
  • July 9: OpenAI ships GPT-5.6 Sol, Terra, and Luna.
  • July 16: Moonshot ships Kimi K3. Much smarter. Also $3.00 in and $15.00 out, about 3x K2.6's input price and 3.75x its output price.
  • September 22: OpenAI ships GPT-6 Sol and Luna at about half the price of their GPT-5.6 versions. The same day, Anthropic ships Claude Opus 5.5, which it says costs 40% less to run than Opus 5.
  • September 25: Fireworks retires K2.6 from its serverless API.
  • September 26: We switch our internal default model from Kimi K3 to GPT-6 Sol.

That's 6 big moves in 5 months. My brokerage account should be so lively.

The board today

MODEL             IN $/1M   OUT $/1M   INDEX   COST TO RUN INDEX
Claude Opus 5.5     4.00      20.00      58        $8,708
GPT-6 Sol           2.00      10.00      48        $1,546
Kimi K3             3.00      15.00      44        $3,658
GPT-6 Luna          0.10       0.50      37          $122
Kimi K2.6           0.95       4.00      27        $1,528

List prices per 1 million tokens. Index score and cost to run from the Artificial Analysis Intelligence Index, at each model's highest reasoning setting. K2.6 is priced at Moonshot's list rate.

How I read the board:

  • Claude Opus 5.5 is the blue chip. Top score, top price. For steps that need real judgment, it earns the premium.
  • GPT-6 Sol beats K3 by 4 points and cost 58% less to run the same suite. That's why we swapped our default.
  • GPT-6 Luna beats our old favorite K2.6 by 10 points and cost 92% less to run. Our value stock got out-valued.

Here's the catch. The board can't tell you which model fits your job.

What "good enough" means on a real job

This is where the auction metaphor earns its keep. In a reverse auction, sellers bid the price down and the lowest bid wins, as long as the goods meet spec. I treat the spec as a reserve price on quality. A model that can't clear it doesn't get a paddle.

Setting the reserve is the hard part. Here's a real one.

A care management customer asked us to screen a patient population of more than 300,000 for chronic care program eligibility. Every decision has to be right, consistent, and explainable to a care team and an auditor. An eligibility call that changes depending on which model ran that day is a liability.

Our first version gave one large model the whole job: read the chart, apply the program rules, score, tier, and submit. It worked. It was also expensive, slow, and hard to audit at 300,000-patient scale.

So we rebuilt it around what each piece does best. The customer's program rules (eligibility, scoring, and tiering) now run in a policy engine they can read, test, and sign off on. The AI agent runs everything around it: it works through the customer's systems to pull each chart, drives the engine, works through errors, and submits the results back. The AI handles the chaos. The engine supplies the structure.

That design cut cost about 11x and made runs about 7x faster. The tier matched our baseline on 26 of 27 test patients, with zero wrong eligibility calls.

Then came my favorite test of the summer. We swapped Kimi K2.6 for Claude Haiku 4.5 and ran the same 40 patients. The submitted decisions were byte-for-byte identical, 40 out of 40. That's the result a compliance team wants to see: the same chart gets the same decision, whichever model runs the workflow. Every decision traces to a rule the customer approved. And it frees us to shop for the model on price and reliability, without re-validating the clinical logic each time.

So the reserve for this auction became reliability at scale: pull every requested chart, drop nothing, invent nothing, and submit clean results, hundreds of patients at a time. That's a harder spec than it sounds, and it turned into its own auction:

  • Kimi K2.6 fit the job, but our host couldn't give us the capacity for a load this size. Out.
  • Claude Haiku 4.5 got creative with the submission format. The customer's API rejected it, and the model looped on shell commands trying to fix things (11 calls per run on average, 36 at worst). Wrong fit for this job. Out.
  • GPT-5.6 Luna was 5x to 9x cheaper than the other options. It also dropped patients. 2 full runs on 574 patients covered 88% and 86%. In 1 test, it invented patient IDs that were never in the request.
  • GPT-5.6 Terra covered 94.8% on a run where Luna covered 69.5%. Better, and still short of the bar.

No bidder cleared the reserve.

So we changed the design again. We started moving the mechanical steps (tracking the patient list, fetching in batches, submitting results) into code, so the agent can't lose or invent a patient. When code owns the patient list, there's no list to invent. After we trimmed the model's job down to 2 tools, a 588-patient run on Luna delivered about 97%. The remaining misses traced to the fetch step, which is the next piece moving into code.

The lesson: the cheapest model that clears the bar depends on how well you design the job around it. Put the approved rules in code and the workflow in the agent, and more bidders clear the reserve. My colleague Weiran Yao calls this "right model, right harness, right job." From the field, I'd add that a good harness is how you get more bidders over the reserve.

When the expensive bid loses anyway

Price and quality don't always move together. Our χ-BENCH team spent over $500 running Claude Fable 5 across 75 healthcare workflow tasks. It passed 24.0% on the first try. Claude Opus 4.6 passed 28.0% and cost less. Fable worked out to $29.20 per passed task, the weakest return in that comparison.

A bigger model bill can buy a worse agent. You only find out by running your own tasks.

The flip side is real too. Our platform groups models into Lite, Regular, and Heavy tiers, and the Heavy tier runs on Claude Opus. For steps that need judgment across long, messy records, that's where the blue chip pays its dividend. When the market moves, we update the tiers. For production agents, we pin the exact model version and re-run the customer's test set before anything changes.

What a re-run week looks like

Here's last week in practice. On September 22, 2 strong new bidders showed up at once. By September 26, our CTO, Frank Wang, had moved our internal default from K3 to GPT-6 Sol, because Sol scored higher on the index and cost less than half as much to run. On September 28, the retired K2.6 started throwing errors, and our engineers had a hotfix moving the same morning.

That week is the whole argument for a gate. Every model change goes through the same steps as a code change: re-run the test set, review what changed, then promote. (It's a boring gate by design. Boring is what you want anywhere near patient data.)

The gate is cheap, too. A test run on a Lite-tier model costs a sliver of what one quarter on the wrong model does. And the week the market moves is the week the savings show up.

How I run the auction for customers

Weiran's post lays out the framework: cost, quality, and compliance. Here's what it looks like on the ground, in the order I do it.

  1. Write the reserve with the customer. What does a correct output look like, field by field? For enrollment, it was the right program, the right tier, every requested patient, and nothing extra.
  2. Build the test set from their data. Real, de-identified cases, including the ugly ones. Public benchmarks are analyst ratings written for someone else's portfolio.
  3. Design the job before you pick the model. Rules the customer has to sign off on go in code. The agent runs the workflow around them. That brings more bidders into the room.
  4. Score cost per completed task. Count retries, dropped records, and human cleanup. A cheap model that drops 14% of patients costs you the 14%, plus the person who has to find them.
  5. Check the paperwork. The model has to run under the customer's BAA, in the right region, with capacity for their volume. K2.6 lost our enrollment auction on capacity alone.
  6. Pin it, then re-run when the market moves. New model, new price, or a retirement notice all trigger a re-run. September gave us 3 of those in 4 days.

Going once, going twice

Our customers get to ignore the ticker. What they see is agents that keep clearing their bar while the cost per task trends down.

Keeping it that way is my team's job. You'll find us on the trading floor. (It's mostly Slack.)

Curious what the winning bid looks like for your workflow? Talk to the ACTAVA.ai team.

Sources

  1. Moonshot AI (Kimi), API pricing: kimi-k3 and kimi-k2.6
  2. Price Per Token, Kimi K2.6 (release date)
  3. OpenRouter, Kimi K3 (release date)
  4. DataNorth, OpenAI releases GPT-5.6 Sol, Terra and Luna
  5. OpenAI, Introducing GPT-6 Sol and Luna
  6. DataNorth, GPT-6 Sol and Luna: OpenAI halves model prices
  7. Anthropic, Claude Opus 5.5
  8. Artificial Analysis, Claude Opus 5.5 vs GPT-6 Sol
  9. Artificial Analysis, Kimi K3 vs GPT-6 Sol
  10. Artificial Analysis, GPT-6 Luna vs Kimi K2.6
  11. Weiran Yao, Right Model, Right Job: Using Cost, Quality, and Compliance to Pick a Model (ACTAVA.ai, Aug 2026)
  12. ACTAVA.ai, χ-BENCH (arXiv 2605.16679)

Joon Lee

Written by

Joon Lee

Head of Forward Deployed Engineering

Forward deployed engineering leader who takes AI agents from prototype to production in regulated healthcare workflows. Joon is Actava.ai's first forward deployed engineer, building and deploying agents across prior authorization, care management, and payer operations. Over 13 years at Salesforce, he helped build three generations of its conversational AI, from NLU-driven chatbots through generative copilots to autonomous agents. As a lead developer on Agentforce, he built core backend services for agent execution, including the graph runtime that adds deterministic control to LLM agents, with a focus on the reliability and security enterprise deployments require. In Salesforce AI Research, he helped lead a small team building AI incubation projects for enterprise customers including Kaiser Permanente, RBC Royal Bank, Equinox, and Adecco, where he ran head-to-head evaluations of agentic reasoning approaches, LLMs, embedding models, rerankers, and RAG architectures. His earlier Salesforce work spanned search relevance, Knowledge search features, and healthcare data models. Before Salesforce, he was an early engineer at Fitbit. He holds a B.A. in Applied Mathematics and Computer Science from UC Berkeley.

Share this