📅 October 7, 2026 · ⏱️ Read time: 5 min · 🔗 Issue No. 29

You're in the loop — Google just priced a frontier-class model at two dollars a million tokens, the kind of number that quietly reorders everyone's budget. Meanwhile, in the week's oddest twist, the internet's most-subscribed human says OpenAI banned him twice for training his own AI at home — a sign the frontier and the fringe are now fighting over the exact same turf.

Today: a 60-second way to bolt a cheap "decision" model onto any workflow, five fresh tools, a copy-paste "second opinion" prompt, and the posts lighting up the timeline.

🔁 The Loop

The moves reshaping AI right now

A new frontier model lands at a fraction of the usual price — and the whole market feels the pull. Illustration: InTheLoop.

  1. Google drops Gemini 4 Argon — and the price of the frontier with it. Google DeepMind's new flagship, Gemini 4 Argon, arrived at an introductory $2 per million input tokens and $10 output (list price later rises to $4/$20), with a 95% discount on cached input — startlingly cheap for a frontier model. It posts strong agentic-coding marks (77.9% on DeepSWE v1.1 and 57.4% on Terminal-Bench 4.0), even as independent evaluators clock its blended intelligence just behind Claude Opus 5.5. The real question is whether Google can actually catch the frontier leaders — watch for OpenAI and Anthropic to answer with price cuts of their own inside two weeks. See the benchmarks.

  2. Etched is fielding chip bids at up to $50 billion. AI-silicon startup Etched is reviewing early investor offers that would value it at $40–50 billion — weeks after a $700M round set it at $21B, which had already doubled a $10.3B mark from July. Etched builds transformer-specialized hardware it claims runs more tokens, faster and cheaper than Nvidia's GPUs, and had booked $1B in orders (including from Jane Street) by mid-year. Read the details.

  3. Anthropic is lining up a $2 trillion IPO. Anthropic is targeting a November listing — possibly before Thanksgiving — at a valuation as high as $2 trillion, with an investor meeting slated for Oct 14 at its San Francisco headquarters. That would rank among the largest tech IPOs ever, even as reporting pegs the company's annual losses near $42B. Here's what's known.

Unlocked reads: a clear-eyed take on whether Argon really closes Google's gap with OpenAI and Anthropic, and what a record quarter of billion-dollar AI rounds says about the froth.

🌊 Deep Current

The quiet rise of the "decision model"

A different kind of model: it doesn't chat, it just decides — fast and cheap. Illustration: InTheLoop.

The new small thing. While everyone watched the frontier labs trade blows, a quieter category shipped all week: tiny, cheap models that don't converse — they decide. Perplexity open-sourced its 27B Decider at $0.04 per million input tokens, and Cloudflare, Amazon and startup Fastino pushed their own in the same stretch.

Why it matters. A decision model takes your app's state, a written question and a fixed set of allowed answers, then returns a calibrated yes/no, a choice or a score — often in tens of milliseconds for a fraction of a cent. For the routing, moderation and "should-the-agent-do-this" calls buried inside every product, that's faster and far cheaper than firing a general chatbot at the problem.

❝

The next wave of AI spend may not be bigger brains at all — it's thousands of tiny, confident judgments, priced by the decision instead of the token.

The catch. Calibration is everything: a fast "no" that's wrong at scale is worse than a slow but careful chatbot. New leaderboards like the Decision Index and JevBench are racing to measure confidence and accuracy, not just latency — and scores still swing widely, with Fastino's GLiDE at 64.81 against rival Jev's 57.91 on one recent panel.

The bottom line. If last year was one giant model doing everything, this year is shaping up around a swarm of specialists. Expect the big labs to start bundling decision endpoints of their own — and expect your stack to get noticeably cheaper the moment you stop paying flagship prices to answer yes or no.

🛠️ The Workbench

Bolt a cheap "decision" model onto your workflow

You're probably paying a flagship model to make tiny routing and triage calls it's wildly overqualified for. Here's how to peel those off onto a decision endpoint in an afternoon.

  1. Find one yes/no or multiple-choice call in your product — say, "is this support ticket urgent?"

  2. Define a fixed answer space: list the allowed outputs explicitly (e.g. urgent / normal / spam).

  3. Pick an endpoint — Perplexity's Decider or Cloudflare Clef for hosted, or a local Strands Decider 2B when the data can't leave your box.

  4. Send the state, the question and the allowed answers, and require a probability alongside each answer.

  5. Set a confidence threshold: below it, escalate to a human or to your bigger model instead of guessing.

  6. Log every decision and spot-check calibration weekly — if accuracy drifts, swap the model, not your whole pipeline.

Sample Prompt: "You are a decision engine, not a chatbot. Given the ticket text and the allowed labels [urgent, normal, spam], return only the single best label plus a confidence from 0 to 1. Ticket: {{ticket_text}}"

🗣️ Overheard

What the timeline's buzzing about

  • 🤖 Atlas grows hands: Boston Dynamics gave its humanoid a new four-finger, 13-degree-of-freedom hand with pressure sensors — and dropped the pinky to get there.

  • 🎭 The video Turing test: Tavus says its real-time avatar Griffin fooled 48% of testers into thinking it was a human on one-minute video calls.

  • 🚫 Banned for building: PewDiePie says OpenAI suspended him twice while he trained a local, uncensored model called Ajax.

  • 📉 Pro gets leaner: OpenAI is trimming ChatGPT Pro's Codex allowance from 20× to 10× the Plus tier starting Oct 30, and power users are grumbling.

  • 🎙️ STT crown: Microsoft's new MAI-Transcribe-2-Streaming is topping real-time speech benchmarks at a 2.5% word-error rate.

🔎 Fresh Finds

Five tools worth a look

  • 🧩 Capy: an "IDE for the parallel age" that runs a fleet of coding agents side by side.

  • 🎨 Comfy Agent: an agent that builds, runs and debugs ComfyUI graphs right on your canvas.

  • 🥒 Pickle: an AI "body double" that can sit in your video calls looking and sounding like you.

  • 📚 Teach Me Anything: turns any topic into a structured, self-paced AI course.

  • 📊 benchlm.ai: a live Decision Index leaderboard for the new wave of decision models.

★ = sponsored placement, if any.

🧪 Prompt Lab

The "second opinion" prompt

Paste this after any model gives you an answer you're about to act on. It forces the model to argue against its own work before you commit.

❝

You just gave me the answer above. Now switch roles and act as a skeptical reviewer who is paid to find what's wrong with it. 1. List the 3 weakest assumptions in your answer and why each could be false. 2. Name 1 scenario where following your advice would backfire. 3. State what single piece of information, if I gave it to you, would most change your conclusion. 4. Give a revised answer only if the review actually changed your mind; otherwise say "no change" and explain why you're confident. Be concrete. No hedging, no flattery.

Want a header image for your own write-up of it? Try this, in a modern gouache style:

❝

A modern gouache illustration: two identical robots seated across a small table, one offering a glowing idea, the other holding up a magnifying glass to inspect it, warm indigo-and-amber palette, soft paper texture, generous negative space, calm and thoughtful mood. No text.

⏪ Rewind

In case you missed it: readers couldn't stop clicking the breakdown of Google's surprise $2/$10 pricing on Gemini 4 Argon — and what it means for everyone else's margins.

Stay in the loop — the InTheLoop team