The pattern this week is price, not capability. Anthropic and OpenAI shipped cheaper models on the same day, 2026-09-22, and the infrastructure stories are about reclaiming what you already pay for.

AI & machine learning

Anthropic releases Claude Opus 5.5. Announced 2026-09-22 at $4/$20 per million input/output tokens, 20% below Opus 5, with cache reads at $0.20 (60% below). Anthropic reports 66.4% on Terminal-Bench 4.0 versus 57.9% for GPT-6 Astra at about 40% of the cost. Vendor numbers, and Anthropic itself says the gap to Fable 5.1 is narrower in daily use. Source

OpenAI ships GPT-6 Sol and Luna at half price. Same day, gpt-6-sol drops to $2/$10 and gpt-6-luna to $0.10/$0.50 per million tokens, 50% below GPT-5.6 promotional pricing. OpenAI reports Sol at 68.8% on DeepSWE v1.1, 1.1 points behind Claude Fable 5, at about 80% lower cost per task. Neither is in Chat yet. Source

Site reliability engineering

Cloudflare Containers leaked disk blocks across tenants. A 2026-09-24 write-up covers a bug reported 2026-09-04: dm-thin pools had skip_block_zeroing set, so a 4 KiB write into a reused 64 KiB block left 60 KiB of another customer’s data readable via /dev/vdc. Researchers found residual data on 20 of 22 nodes. Cloudflare removed the flag, then retired every running disk and cached image snapshot, since zeroing only covers new allocations. Source

Cloudflare cut 100 TB of RAM by shrinking a consistent-hash ring. Pingora Backend Router used up to 6 GB per process on pingora-ketama points. Packing the {hash: u32, index: u32} struct into 6 bytes saved 25%; deriving the error formula for k hashes per server showed 90% of the points added nothing measurable. Source

Observability

Prometheus 3.15.0 ships OpenMetrics 2.0 scraping. Released 2026-09-24, it also adds Unix-socket scrape targets, zstd scrape responses behind zstd-scrape, runtime GOMEMLIMIT re-detection, and log level changes on reload via runtime.log_level; --log.level is deprecated. XOR2 float chunk encoding is stable, but check that Thanos sidecars and anything else reading TSDB directly support it first. Source

OTel and Prometheus interoperability survey: friction down, hybrids everywhere. Published 2026-09-22 from 81 screened end users: the share calling the two hard to use together fell from 29% in 2024 to 10%. For infrastructure metrics, 72% use Prometheus exporters, 57% use OTel receivers, and nearly half run both. Source

  • Honeycomb documents its adaptive_tail_sampling Collector processor: first-match rules and thresholds in ot=th TraceState. Source
  • Kubernetes 1.37.1 fixes a DeviceTaintRule with no deviceSelector matching every DRA device cluster-wide. Source
  • Grafana 13.3 imports Prometheus and Mimir Alertmanager configuration into Grafana Alerting (preview). Source
  • Lorin Hochstein on GitHub’s 2026-09-13 incident: a cleanup job paced by replica lag saturated the primary instead. Source

My take

The Cloudflare Containers fix is the one to copy. Flipping skip_block_zeroing fixed new allocations, but already-mapped blocks in running disks and image caches stayed dirty, so they rebuilt the fleet. If your mitigation only changes a default, ask what existing state still carries the old behavior.

See you next week.