Daily Digest - 2026-09-20
Shopify founder Tobi Lütke argues that MCP-versus-CLI is the wrong frame. What matters is whether agents can work in persistent execution environments.
Shopify founder Tobi Lütke argues that MCP-versus-CLI is the wrong frame. What matters is whether agents can work in persistent execution environments.
Looking at more than 43,000 model calls shows that fragmented and cut-down inference setups, not changed weights, are why production models lag behind benchmark claims.
Ion Stoica explains why autonomous coding agents learn to game benchmarks and fail in production when written prompts and test environments do not match real-world requirements.
Custom Jev-style models using parallel constrained decoding could make existing agent workflows far more token-efficient. Consider an agent workflow where an LLM reviews every support ticket, invoice, or claim before the next step.
Surge AI shows how training Kimi K2. 7 entirely with reinforcement learning across 1,700 coding tasks improves efficiency and transfers across benchmarks without using supervised fine-tuning.
How OpenAI shifted toward an autonomous software factory model driven largely by non-engineers. Non-engineering teams at OpenAI, including legal, recruiting, and finance, now rely on Codex and ChatGPT Work as their primary day-to-day tools.
How to turn custom customer projects into core product features rather than one-off consulting jobs. Forward-deployed engineers are not just technical consultants.
See how giving coding agents domain-specific evaluation skills lets them audit, diagnose, and benchmark AI applications on their own.