The pattern this week is price, not capability. Anthropic and OpenAI shipped cheaper models on the same day, 2026-09-22, and the infrastructure stories are about reclaiming what you already pay for.
AI & machine learning
Anthropic releases Claude Opus 5.5. Announced 2026-09-22 at $4/$20 per million input/output tokens, 20% below Opus 5, with cache reads at $0.20 (60% below). Anthropic reports 66.4% on Terminal-Bench 4.0 versus 57.9% for GPT-6 Astra at about 40% of the cost. Vendor numbers, and Anthropic itself says the gap to Fable 5.1 is narrower in daily use. Source
OpenAI ships GPT-6 Sol and Luna at half price. Same day, gpt-6-sol drops to $2/$10 and gpt-6-luna to $0.10/$0.50 per million tokens, 50% below GPT-5.6 promotional pricing. OpenAI reports Sol at 68.8% on DeepSWE v1.1, 1.1 points behind Claude Fable 5, at about 80% lower cost per task. Neither is in Chat yet. Source
Site reliability engineering
Cloudflare Containers leaked disk blocks across tenants. A 2026-09-24 write-up covers a bug reported 2026-09-04: dm-thin pools had skip_block_zeroing set, so a 4 KiB write into a reused 64 KiB block left 60 KiB of another customer’s data readable via /dev/vdc. Researchers found residual data on 20 of 22 nodes. Cloudflare removed the flag, then retired every running disk and cached image snapshot, since zeroing only covers new allocations. Source
Cloudflare cut 100 TB of RAM by shrinking a consistent-hash ring. Pingora Backend Router used up to 6 GB per process on pingora-ketama points. Packing the {hash: u32, index: u32} struct into 6 bytes saved 25%; deriving the error formula for k hashes per server showed 90% of the points added nothing measurable. Source
Observability
Prometheus 3.15.0 ships OpenMetrics 2.0 scraping. Released 2026-09-24, it also adds Unix-socket scrape targets, zstd scrape responses behind zstd-scrape, runtime GOMEMLIMIT re-detection, and log level changes on reload via runtime.log_level; --log.level is deprecated. XOR2 float chunk encoding is stable, but check that Thanos sidecars and anything else reading TSDB directly support it first. Source
OTel and Prometheus interoperability survey: friction down, hybrids everywhere. Published 2026-09-22 from 81 screened end users: the share calling the two hard to use together fell from 29% in 2024 to 10%. For infrastructure metrics, 72% use Prometheus exporters, 57% use OTel receivers, and nearly half run both. Source
Quick links
- Honeycomb documents its
adaptive_tail_samplingCollector processor: first-match rules and thresholds inot=thTraceState. Source - Kubernetes 1.37.1 fixes a
DeviceTaintRulewith nodeviceSelectormatching every DRA device cluster-wide. Source - Grafana 13.3 imports Prometheus and Mimir Alertmanager configuration into Grafana Alerting (preview). Source
- Lorin Hochstein on GitHub’s 2026-09-13 incident: a cleanup job paced by replica lag saturated the primary instead. Source
My take
The Cloudflare Containers fix is the one to copy. Flipping skip_block_zeroing fixed new allocations, but already-mapped blocks in running disks and image caches stayed dirty, so they rebuilt the fleet. If your mitigation only changes a default, ask what existing state still carries the old behavior.
See you next week.