Time to First Token
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization
About Time to First Token
Time to First Token is a structured 10 week curriculum covering LLM inference serving and optimization, designed around 30 minutes of study a day. It walks through batching, KV caching, quantization, throughput versus latency tradeoffs and the serving stack choices that decide your infrastructure bill. Worth the hour a week the moment self-hosting or inference cost stops being a rounding error.
More Guides
.png)
UpGuard: 16,326 Supabase Databases Exposing Data
UpGuard Research found 16,326 Supabase databases readable without authentication, more than half showing personal information, from apps built all over the world. Published 25 September 2026, it is the largest study of its kind. Why it matters for builders: Supabase is the default backend for vibe-coded apps, and tables created through SQL, migrations or the API do not get row-level security switched on. The agent scaffolds the table; nobody writes the policy. The anon key then reads everything. Read the study, then query every table in your project with the anon key before you share a link. Our companion checklist walks through the fix.

AI Coding Made CI a Bottleneck, So Linear Reworked Theirs
Agents made shipping code faster; validating it did not keep up. Linear’s test suites almost quadrupled in 2026, yet the team brought pull-request wait time down from over six minutes to just over five and cut runner time per test roughly in half. The four moves: upgraded infrastructure and tooling (faster third-party runners instead of GitHub Actions), optimised the jobs that gate other work, reduced repeated setup, and made test execution more efficient. The codebase is TypeScript, but most of it transfers. Read it if your agents now open more PRs than your CI can absorb, which is the new normal for a small team running Claude Code or Codex all day.

Jev in 25 Lines of Python
NobodyWho strips the mystique from decision models. Load any GGUF model with llama-cpp-python, give it lettered choices in the prompt, read the logits of the next token for each letter, and normalise them into probabilities. That is the whole trick. Why it matters: If you pay an LLM to classify support tickets, route requests or score leads, this shows the mechanic underneath the hosted decision-model APIs: a single forward pass, no generation, no fine-tuning, and the data never leaves your machine. The post is honest about the limits too: the probabilities are not calibrated the way a trained decision model’s are. Read it before you pick between a hosted API, a local model or 25 lines of your own.

AI Overviews Cut CTR by 23.1% in France
Google launched AI Overviews in France on July 22, 2026. Ahrefs tracked 963 domains in Google Search Console for 28 days before and 9 days after, and the most exposed sites lost 23.1% of their click-through rate. If search is your main acquisition channel, this is the data to read before you plan next quarter: it shows which kinds of pages get answered without the click. Read the study on the Ahrefs blog.

We Pinned Our Model Version. The Provider Deprecated It Anyway.
Pinning a model version does not protect you from deprecation. This Towards Data Science piece argues that the recurring cost of production AI is not inference but re-qualification: the eval reruns, prompt retuning and regression testing you owe every time a model changes under you. It breaks down what that tax covers and how to budget for it before a deprecation notice forces the work on you. Read the guide on Towards Data Science.