AI Engineering Signal #77
Grok 4.5 ships, ranks fourth on Artificial Analysis Intelligence Index at half frontier cost
Signals
Grok 4.5 ships, ranks fourth on Artificial Analysis Intelligence Index at half frontier cost
reprice agentic routing assumptions and run head-to-head evals before next sprint.
TechCrunch
Anthropic's Fable 5 orchestrator pattern cuts cost nearly in half
cheap-model-under-orchestrator is now a benchmarked production pattern, not a theory.
Cognition SWE-1.7 reaches near GPT-5 and Claude Opus 4.x coding performance
coding agent procurement decisions from six months ago may now be mispriced; rerun evals.
Web
Databricks benchmarks coding agents on a multi-million-line codebase
enterprise codebases expose agent failure modes that standard benchmarks miss; audit your eval harness.
Web
OpenAI publishes methodology for separating signal from noise in coding evals
benchmark inflation is now documented by a lab; cross-validate before trusting any single-benchmark claim.
Web
Mistral releases Robostral Navigate robotics navigation model
open-weight nav baseline now exists; proprietary robotics nav stack procurement faces direct cost pressure.
Web
The Take
Frontier capability is compressing from both ends: new models hit near-top-tier performance at lower cost while orchestration patterns and eval methodology are mature enough to make cost-per-task the operative metric. Teams routing all traffic to a single flagship model are overpaying and flying blind on quality.
Subscribe
Related Signals