Blog

Daily posts on AI infrastructure, LLMs, GenAI, agents, recommendation systems, and ML engineering — with diagrams and runnable notebooks.

2026-10-04

95% Harmful, Zero Red Flags: The Agent Handoff Problem Nobody Tests

Agents / AI safety (RogueHandoff-20 — multi-agent handoff injection)
2026-10-03

100,000 Videos in Your Pocket: How Douyin's Recommender Reads Your Whole Life in O(1)

Recommendation systems (ultra-long sequence modeling)
2026-09-30

The Model That Failed Its Own Safety Test: Why OpenAI Shelved GPT-6.1 Astra

Tech news / AI safety (GPT-6.1 Astra shelved; agent alignment evals)
2026-09-29

Tokens Are a Crutch: Why Byte Models Win the Long Game

NLP updates / ML architectures (byte-level models — "Breaking the Token Ceiling", arXiv 2609.12303)
2026-09-28

5.7x, 512 GPUs, One Endpoint Across the Pacific: AI's Report Card Just Grew Up

AI infrastructure / benchmarks (MLPerf Inference v6.1)
2026-09-27

Thinking Twice Can Make You Dumber: The Test-Time Compute Playbook Behind September's Smartest Models

AI infrastructure / ML architectures (test-time compute scaling)
2026-09-25

The Paper That Answers Back: Stanford Turned 100 Research Papers into Agents

Agents / AI for science (Paper2Agent — papers as MCP agents)
2026-09-24

The Model That Doesn't Talk

Tech news / ML architectures (TypeSafe Jev — decision-output transformer)
2026-09-23

The Tuesday Intelligence Went on Sale

Tech news / LLM economics (frontier price war)
2026-09-22

One Update, One Quarter-Turn: The Attention Layer That Learned to Rotate

ML architectures (linear attention / Kimi Delta Attention)
2026-09-21

1,107 Tokens Per Second: The LLM That Doesn't Type

GenAI + ML architectures (diffusion LLMs)
2026-09-20

The Two-Phase Machine: Your LLM Request Is Two Jobs in a Trench Coat

AI infrastructure (LLM inference serving)
2026-09-19

One Model, Three Jobs, Zero Retraining: The Spreadsheet Just Got Its Foundation Model

Tech news / ML architectures (tabular foundation models)