AI Engineering Signal #94
GPT-5.6 price cut driven by recursive self-optimization
Signals
GPT-5.6 price cut driven by recursive self-optimization
rerun cost models now; equivalent intelligence dropped 13x in four months, making Q1 budget assumptions stale.
Latent Space
Qwen3.8-27B fits in 17GB VRAM at strong coding benchmarks
direct swap candidate for local inference rigs running larger, costlier models.
Web
EU AI Act disclosure rules took effect Sunday
audit every deployed chatbot and AI-generated media pipeline for human-disclosure logic now.
Web
llama.cpp adds MTP and DeepSeek V4 Flash support
DeepSeek V4 Flash inference now routes through llama.cpp without a custom serving layer.
GitHub
Quanta analysis questions whether AI reasoning is right for wrong reasons
benchmark pass rates are insufficient; add adversarial probing before shipping reasoning agents.
Web
OpenAI, Anthropic, Google DeepMind, Meta co-sign letter to pace AI development
watch for self-imposed capability gates that could affect API roadmap timing.
Latent Space
The Take
Inference costs are collapsing faster than procurement cycles can track, while the same frontier labs shipping those cuts are now signaling for a slowdown — your cost models, capability assumptions, and compliance posture are all expiring on different schedules simultaneously.
Subscribe
Related Signals