Inference Cost
Curated collection of thoughts and builds centered around Inference Cost.
Daily Briefings 20
- AI infrastructure strain as data center power demand surges; token costs fall 41% as enterprises shift models
- Reducto ships r-1 document parser at 1¢/page; OpenBMB releases MiniCPM5-2B small language model
- Google's Gemini cuts video token usage 88%; Anthropic releases Fable 5.1 with 45% cost reduction
- Google's Gemini Omni 1.1 Flash extends video generation; agent sandbox pricing comparison emerges
- FreeToken runs 753B MoE models on single GPU; agentic token consumption surges 14x
- Alibaba Qwen 3.8 open-weights release; GLM-5.3 claims strongest coding model via post-training
- Google Gemini 3.7 Flash undercuts predecessor 50%; DeepSeek open-sources agent harness
- SpaceXAI's Grok 4.6 matches frontier performance at lower cost; Dyna-2 scales robot learning to 1M video hours
- NVIDIA Releases Nemotron 3.5 Lightning MoE and LTX-2.5 Open Video Model
- Claude Code Auto Mode becomes default; agent energy consumption tracked at 600x chat baseline
- Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model
- DeepSeek V4 Flash Update Matches GPT-5.6 Luna at 60% Lower Cost; Thinking Machines Releases Smaller Inkling Model
- OpenAI cuts GPT-5.6 Luna pricing 80%; Google DeepMind releases Gemini Robotics 2 models
- Token Saver MCP cuts PDF costs 90-99%; Moonshot open-sources MoonEP for MoE training
- Cursor's Agent Swarm Achieves 100% SQLite Rebuild; Black Forest Labs Releases FLUX 3 Multimodal Model
- Anthropic releases Claude Opus 5 at unchanged pricing; Datalab Marker v2 benchmarks document processing
- Gigatoken tokenizer hits 24.5 GB/s; Cursor Router cuts inference costs 30–50%
- OpenAI models breach Hugging Face during internal security test; Google releases three new Gemini Flash variants
- Anthropic cuts Claude Fable limits; GPT-5.6 deletes files in full access mode
- GPT-5.6 Sol reported near Fable 5 at a third the cost; Kyutai ships open-weight MuScriptor