Who Actually Did the Work?
TL;DR — Key Takeaways
- Anthropic and OpenAI cut frontier model prices within 90 minutes; Opus 5.5 runs 40% cheaper.
- Six global banks published principles for AI agents that shop and pay on a customer's behalf.
- Contractors grading ChatGPT were removed for using AI to do the grading.
Anthropic and OpenAI cut prices on their best models about 90 minutes apart. Six global banks published the rules they want before AI agents start shopping with your customers' cards. And contractors hired to grade ChatGPT lost the work for using AI to do it. AI got cheaper to hire this week. It also got harder to supervise.
The Bottom Line
All three stories ask one question: who did the work, and who answers for it? The model that reads your invoices now costs pennies. The buyer at your checkout may be software. And the teammate who finished early may have had help. So price the automation, write the rule, keep the receipt. Tell me in the comments: would you fire someone for using AI on work you hired them to do by hand? Or is the missing rule the real failure?
Top 3 Stories This Week
01 · Anthropic · OpenAI · Model Pricing
The Best AI Got Cheaper Twice in One Afternoon
On September 22, Anthropic released Claude Opus 5.5. Roughly 90 minutes later, OpenAI answered with GPT-6 Sol and GPT-6 Luna. Both companies led with the same word: cheaper. Anthropic says Opus 5.5 matches its top model, Fable 5.1, on most work and costs 40% less to run than Opus 5, at $4 per million input tokens and $20 per million output (a token is about three-quarters of a word). OpenAI cut Sol to $2 and $10, half its predecessor, and priced Luna, built for clerical work like summarizing and pulling data out of documents, at $0.10 and $0.50. The test owners should watch is Zapier's AutomationBench, which runs real business workflows across connected apps. Opus 5.5 scored 40.0%. OpenAI says Sol finishes a task there for 27 cents. Read that twice. The best models still fail most multi-app workflows. What changed is the price of trying. The automation you shelved in spring costs a fraction today.
Microservice Opportunity: Re-quote every automation you shelved on cost. Then build a router: a microservice (one small AI tool doing one specific job) that sends high-volume chores such as sorting email or pulling invoice fields to the cheapest model, sends judgment calls to Opus or Sol, and hands anything the model flags as uncertain to a person. Log cost per task from day one. When prices drop again next quarter, you switch models in an afternoon.
02 · NatWest · Bank of America · Agentic Commerce
Six Banks Want Rules Before AI Agents Go Shopping
Picture an order that lands at 2 a.m., placed by software instead of a person. If it is wrong, who pays? On September 22, six banks (NatWest Group, ASB Bank, Bank of America, Capital One, Commonwealth Bank of Australia and ING Group) published shared principles for agentic commerce, the term for AI agents that find, choose and pay for products on a shopper's behalf. The paper, Building Trust in Agentic Commerce, sorts the problem into five areas: transparency, safety, privacy and data, choice, and interoperability. It warns that scams, fraud and disputes could all rise. It also flags a quieter risk: an agent may rank products by which merchant pays the biggest commission, not by what fits the buyer. Bank of America's head of global payments solutions named the open questions plainly: identity, authorization, fraud prevention and liability. The banks are not against agent shopping. They call it a plausible mainstream way to pay, and a second paper on how to apply the principles is coming. For a merchant, that means the rules on who eats a bad agent order are being drafted now, without you.
Microservice Opportunity: Build a checkpoint for orders placed by software. A microservice flags orders that look automated (odd hours, brand-new accounts, checkout completed in seconds), texts the human buyer to confirm anything above a dollar threshold before it ships, and files a dispute packet for every order: the confirmation, the order record, the delivery photo. When the first agent chargeback arrives, you answer with evidence in minutes instead of losing the sale by default.
03 · OpenAI · Mercor · AI at Work
OpenAI's Graders Lost the Job for Using AI
OpenAI sells AI as the fastest way to get work done. On September 22, 404 Media reported that contractors hired to grade ChatGPT's answers were removed for doing exactly that. Internal documents describe more than ten thousand contractors across OpenAI's rating work and bar them from using AI tools, including GPTZero and Grammarly, while they work. Supervisors spot offenders by repeated phrasing, heavy use of em dashes and suspiciously fast finishes. Mercor, the staffing firm that supplies many of the raters, confirmed it removes anyone it confirms used AI on a task. One removed worker summed up the temptation: "I just needed a little boost." The rule makes sense here. OpenAI pays for human judgment, and labs worry that AI grading AI slowly makes models worse. The irony is loud. The lesson is quieter and belongs to every employer: a rule about AI that nobody wrote down is a dismissal waiting to happen. Does your team know where your line sits?
Microservice Opportunity: Write the rule per task, then make it painless to follow. Label each job: AI allowed, AI with disclosure, or human only. A microservice sits where work gets submitted (proposals, estimates, client reports), asks one question, "Did AI draft any of this?", logs the answer, and routes AI-drafted client work to a person for a final read. You stop guessing who wrote what, and nobody loses a job over a rule they never saw.
This week in numbers: 40%: less to run Claude Opus 5.5 than Opus 5. OpenAI halved GPT-6 Sol's price the same afternoon. · 6 banks: drafting the rules for who pays when an AI agent buys the wrong thing from your store. · 10,000+: contractors in OpenAI's grading work, barred from using AI to do it.
Sources
- Anthropic: Introducing Claude Opus 5.5
- OpenAI: Introducing GPT-6 Sol and Luna
- TechCrunch: OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes
- Financial IT: Global Banks Collaborate on Principles for Trusted Agentic Commerce
- Gizmodo: Big Banks Say They're Uneasy About People Shopping via AI Agents
- 404 Media: People Training OpenAI's AI Fired for Using AI to Train the AI
- AI Weekly: OpenAI Fires Contractors Caught Using AI to Grade ChatGPT