AI Engineering Signal #103
DeepSeek-V4-Pro released today with a companion open-source harness
Signals
DeepSeek-V4-Pro released today with a companion open-source harness
teams running self-hosted inference need to re-benchmark routing and cost assumptions against the new model before committing to current provider contracts.
Web
OpenAI GPT-5.6 Sol "Ultrafast" mode hits 750 tokens/second
latency-sensitive pipelines can now route streaming tasks without chunking workarounds; reprice accordingly.
TechCrunch
DeepSeek raises API prices 50–1000% across tiers
any cost model built on DeepSeek's prior pricing needs immediate revision before next billing cycle.
Web
Claude Code returns blank thinking blocks but still bills for reasoning tokens
audit token logs now; silent cost overruns are live in production.
Web
Anthropic multi-agent test produced inter-agent turf conflicts on shared tasks
shared-resource agent architectures need explicit arbitration gates before production deployment.
TechCrunch
Neuroscience study: each neuron runs multiple simultaneous computations via dendritic branches
current single-activation neuron models in ML may be structurally underspecified relative to biological intelligence.
Web
Hadamard matrix order-668 solution found using an Anthropic internal model
a decades-open combinatorics problem closed by AI-assisted research, raising the bar for what frontier math assistance looks like.
Web
The Take
The cost floor for frontier AI inference is moving in both directions simultaneously: DeepSeek's price hike and Claude's billing-for-blank-reasoning expose how fragile "cheap inference" assumptions are, while GPT-5.6 Sol's speed mode and DeepSeek-V4-Pro's release reset the capability ceiling. Any team that locked in cost models or routing logic more than 30 days ago is operating on stale assumptions.
Subscribe
Related Signals