Qwen 3.5 Open Source Shocks, AI Engineering Deepens, METR Evaluates AI Productivity
W03

Qwen 3.5 Open Source Shocks, AI Engineering Deepens, METR Evaluates AI Productivity

Alibaba's Qwen 3.5 releases 8 model sizes (397B to 0.8B), METR studies exponential AI productivity timelines, Dylan Patel analyzes $200B AI capex

721articles
25+sources
Y

In this issue: Qwen 3.5 model family → METR AI productivity evals → Dylan Patel on chip wars → Google AI updates → HuggingFace Modular Diffusers → Code review crisis → References


Qwen 3.5: Open Source AI Milestone (And the Crisis Behind It)

Alibaba's Qwen team densely released 8 model sizes (397B to 0.8B) over the past few weeks, shaking the open-source community. Highlights:

ModelFeature
397B / 235BFlagship reasoning + multimodal
27B / 35BExcellent coding tasks on Mac
2B4.57GB (1.27GB quantized), full reasoning + vision
0.8BIdeal for edge devices

These were achieved with far fewer resources than OpenAI/Google — Simon Willison called it "a remarkable achievement."

However, in the same period, technical lead Junyang Lin abruptly announced his resignation on X, followed by multiple core members. The trigger: Alibaba hired a new manager from Google's Gemini team to take over Qwen. Open source AI's biggest crisis was emerging.


METR: Exponential AI Productivity Timeline Evaluation

METR's Joel Becker shared their evaluation framework on the Latent Space podcast:

  • "Time Horizon Evals": Measuring how long AI can independently complete tasks
  • AI capability ceilings are growing exponentially, not linearly
  • Threat model: When AI can independently handle 6-month tasks, the software industry will fundamentally reorganize
  • Currently, AI can autonomously handle tasks in the 30-60 minute range — and advancing rapidly

"We need reliable evaluation systems in place before AI capabilities exceed human oversight."


Dylan Patel: $200B AI CapEx — Who's Burning, and Where

SemiAnalysis's Dylan Patel broke down AI industry capital allocation on the Latent Space podcast:

  • $200 billion in annual AI infrastructure spend (Google, Meta, Microsoft, Amazon)
  • Chip war core: HBM memory and CoWoS packaging are the bottlenecks; TSMC is the key
  • Google may have zero profits by 2027: AI investment compressing ad margins
  • Blackwell architecture's impact far exceeds Hopper — higher performance, but also more expensive

Google AI February Recap: Nano Banana 2 and AI Mode Canvas

Google summarized February's AI releases, with key highlights:

  • Nano Banana 2: Best image generation model, outperforming Midjourney v7 on specific benchmarks
  • AI Mode Canvas: Document creation and interactive tools directly in Google Search, available to all US users
  • Gemini Translate update: New "understand" and "ask" buttons to navigate linguistic complexity

HuggingFace: Modular Diffusers and GGML Merger

Two major HuggingFace developments this week:

  1. Modular Diffusers launch: Composable diffusion pipeline building blocks — developers can assemble image generation workflows like LEGO
  2. GGML and llama.cpp join HuggingFace: Ensuring the long-term progress of local AI; two most important open-source inference libraries now under one roof

Code Quality's New Challenge: AI Acceleration Creates PR Review Pressure

Latent Space published the controversial "Code Review Is Dead" article, with striking data:

  • Teams with high AI adoption merge 98% more PRs
  • But PR review time increased 91%
  • Humans can no longer read all AI-generated code

Future quality gates lie in Spec-Driven Development: Review intent and specifications, not code itself.


References

Qwen 3.5

METR Evaluation

Dylan Patel

Google AI

HuggingFace

Code Review

AIOpen Source ModelsQwenAI EvaluationDev Tools