AI Engineering Signal #82
Thinking Machines releases Inkling, a 975B-A41B multimodal MoE under Apache 2.0
Signals
Thinking Machines releases Inkling, a 975B-A41B multimodal MoE under Apache 2.0
now ranked top US open-weight model; benchmark against Llama 4 Maverick before next procurement decision.
Web
Claude Opus 4.6 tops LMArena after factuality scoring added
re-run routing logic if you deprioritized Opus 4.6 without factuality benchmarks.
Claude web-fetch prompt-injection exfiltration documented
audit any Claude deployment using web-fetch tools for data exfiltration paths now.
Simon Willison
xAI open-sources Grok Build under Apache 2.0
evaluate this agentic build scaffold for CI/CD agent pipelines this week.
GitHub
Gemma 4 chat template update fixes tool-calling and output laziness
redeploy Gemma 4 tool-use pipelines against updated templates before blaming the model.
Web
Apple in talks to acquire on-device compression startup PrismML
iOS inference assumptions and CoreML API surface may shift; watch before locking deployment targets.
Web
The Take
Open-weight quality is closing on proprietary flagships faster than cost models assumed. Inkling leading US open-weight rankings while Opus 4.6 tops factuality scoring means the "closed for quality, open for cost" heuristic needs a fresh benchmark pass — and the Claude exfiltration finding confirms that expanding agent permissions without an audit is now a documented, not theoretical, risk.
Subscribe
Related Signals