Free the Models: Harness Design at the Frontier
Replit stopped routing work for the model and let the model decide. Here are the numbers.

About Free the Models: Harness Design at the Frontier
Replit argues that model routers have a built-in ceiling: the router is always less capable than the model it is choosing for. So Replit Agent now lets the frontier model decide how hard to think, when to hand work off, and which specialist subagent (code, UI design, slides, writing) takes each step.
Why it matters: on DeepSWE, Replit Agent scored 72% at $2.11 a task against 74% at $4.43 for Astra alone; on Terminal-Bench, 49% at $2.53 against 60% at $5.86. Models also use the same freedom differently: GPT-6 Astra reused briefed workers 42% of the time, while Fable models rarely used general workers.
Read it if you are building an agent and are about to hard-code which model or step handles what.
More Guides

Slopsquatting: Your AI Invented a Package, and an Attacker Was Waiting
AI coding tools sometimes recommend packages that do not exist, and attackers register those names with malware inside. This guide walks through the USENIX Security 2025 study behind the "one in five" figure: 576,000 code samples from 16 models, 19.7% of suggested packages hallucinated, 205,474 unique fake names, and 43% of hallucinations repeating when the same prompt ran ten times. Why it matters: repeatable hallucinations are what make the attack work, because an attacker can predict the name. If your agent can run npm install or pip install on its own, it is the one typing the command. The fixes it lists: check a package’s age, downloads and repository before installing; lockfiles with hashes; npm install --ignore-scripts; reject packages younger than about 90 days; and never let an agent self-install unvetted dependencies.
.png)
UpGuard: 16,326 Supabase Databases Exposing Data
UpGuard Research found 16,326 Supabase databases readable without authentication, more than half showing personal information, from apps built all over the world. Published 25 September 2026, it is the largest study of its kind. Why it matters for builders: Supabase is the default backend for vibe-coded apps, and tables created through SQL, migrations or the API do not get row-level security switched on. The agent scaffolds the table; nobody writes the policy. The anon key then reads everything. Read the study, then query every table in your project with the anon key before you share a link. Our companion checklist walks through the fix.

AI Coding Made CI a Bottleneck, So Linear Reworked Theirs
Agents made shipping code faster; validating it did not keep up. Linear’s test suites almost quadrupled in 2026, yet the team brought pull-request wait time down from over six minutes to just over five and cut runner time per test roughly in half. The four moves: upgraded infrastructure and tooling (faster third-party runners instead of GitHub Actions), optimised the jobs that gate other work, reduced repeated setup, and made test execution more efficient. The codebase is TypeScript, but most of it transfers. Read it if your agents now open more PRs than your CI can absorb, which is the new normal for a small team running Claude Code or Codex all day.

Jev in 25 Lines of Python
NobodyWho strips the mystique from decision models. Load any GGUF model with llama-cpp-python, give it lettered choices in the prompt, read the logits of the next token for each letter, and normalise them into probabilities. That is the whole trick. Why it matters: If you pay an LLM to classify support tickets, route requests or score leads, this shows the mechanic underneath the hosted decision-model APIs: a single forward pass, no generation, no fine-tuning, and the data never leaves your machine. The post is honest about the limits too: the probabilities are not calibrated the way a trained decision model’s are. Read it before you pick between a hosted API, a local model or 25 lines of your own.

AI Overviews Cut CTR by 23.1% in France
Google launched AI Overviews in France on July 22, 2026. Ahrefs tracked 963 domains in Google Search Console for 28 days before and 9 days after, and the most exposed sites lost 23.1% of their click-through rate. If search is your main acquisition channel, this is the data to read before you plan next quarter: it shows which kinds of pages get answered without the click. Read the study on the Ahrefs blog.