📜 The Primary Source

Friday, September 11, 2026

12 stories · Standard format

Generated with AI from public sources. Verify before relying on for decisions.

🎧 Listen to this briefing or subscribe as a podcast →

Agentic cache pricing just collapsed by two orders of magnitude courtesy of a new Chinese model. Elsewhere in this edition: the USPS inspector general confirms a decade of network restructuring has failed to fix agency finances, and two Supreme Court briefs could rewrite the zoning playbook for private religious gatherings nationwide.

Frontier AI (Practitioner)

DeepSeek V4.1-Flash: $0.003/M Cached Input, 552B MoE, 890-Byte KV Cache — An Order-of-Magnitude Repricing of Agentic Inference

DeepSeek released V4.1-Flash on Thursday with off-peak cached-input pricing of $0.003 per million tokens — roughly 100x cheaper than Claude Opus 5's $0.50 and GPT-5.6 Sol's $0.40 — enabled by a novel Causal Encoder-Decoder (CED) architecture that compresses KV cache to 890 bytes per token (one-quarter V4-Flash's size) and activates only 8B parameters during prefill and 16B during decode from a 552B total. On Terminal-Bench 2.1, DeepSeek reports V4.1-Flash at 90.6, ahead of GPT-5.6 Sol (88.8) and Kimi K3 (88.3); on DeepSWE v1.1, the model scores 74.2 versus Opus 5's 74.0. Starting September 14, DeepSeek will automatically route all V4-Pro API requests to V4.1-Flash at the lower price. Peak-hour pricing (01:00–04:00 and 06:00–10:00 UTC Monday–Friday) doubles the rates. The model's MIT-licensed open weights and concurrency limit expansion from 500 to 2,500 requests accompany the launch.

The cost gap here is structural, not marginal. An agent reusing a 500,000-token repository prefix across 100 requests pays roughly $0.15 on V4.1-Flash off-peak versus $20–$25 on Anthropic or OpenAI flagships — a 133–166x difference that no benchmark delta between models can bridge if the workload is input-heavy and repetitive. DeepSeek's benchmark claims (Terminal-Bench 2.1: 90.6, DeepSWE: 74.2) are self-reported and not yet independently reproduced by external labs; the architectural claims (CED, CSA2, Engram lookup table, DSpark speculative decoding) are also unverified. What is verifiable immediately is the price, the MIT license, and the automatic V4-Pro migration starting September 14 — which means any production deployment on the V4-Pro identifier needs regression testing before that date. For agentic workflows where the same long context is read repeatedly, this is the new pricing floor every other provider will have to address.

Verified across 5 sources: VentureBeat · Forkast · DeepSeek · WCCFtech · IndexBox

Anthropic Releases Smart Reports for Enterprise and Fable 5.1 Leads BenchLM Agentic Leaderboard at 80.2

Following the Anthropic platform cost-optimization audit we covered Tuesday — which demonstrated up to 73% cost reductions via prompt-pattern removal — the company's September release notes introduce smart reports (beta) for Enterprise administrators. The tool analyzes team usage to identify friction points and flag repeated tasks worth packaging as shared skills. Separately, after Artificial Analysis tied Fable 5.1 and GPT-6 Astra at 53 points earlier this week, Fable 5.1 now leads the BenchLM agentic leaderboard with an 80.2, ahead of Claude Opus 5 (78.1). Anthropic's own benchmarks show Fable 5.1 pricing reductions of roughly 25% for typical workloads and up to 45% for highly agentic work through lower cache-read rates ($0.25/M).

Smart reports operationalize the exact cost-visibility audit Anthropic recommended earlier this week, moving from reactive terminal checks to proactive identification of token-burning workflows. The BenchLM score confirms Fable 5.1's agentic lead over Opus 5 is real but modest (80.2 vs. 78.1), meaning the routing question is still a per-task decision. Independent testing at The New Stack found Fable 5.1 solved 1 of 5 Terminal-Bench-Science tasks on a $12/task budget versus 0 for Fable 5, suggesting the headline performance gains are genuine but narrower under production constraints than vendor benchmarks imply.

Verified across 4 sources: Anthropic Support · BenchLM · Anthropic · The New Stack

Cognition Releases SWE-2: Kimi K3-Based Coding Agent Scores 50% on FrontierCode 1.1 at 64% Lower Cost Than Fable 5.1, 81% Fewer Turns

Cognition released SWE-2 on Thursday, a specialized coding model built on the Moonshot AI Kimi K3 open-weight base we've been tracking. Trained with a novel reinforcement learning method that jointly trains multiple effort levels in a single run, SWE-2 scores 50.0% on FrontierCode 1.1 Main (per Cognition's self-reporting) at a claimed 64% lower cost than Fable 5.1. SWE-2 medium matches SWE-1.7's score at 58% fewer turns (18 median steps vs. 48) and 81% lower average cost.

The 18-step median to first code edit versus SWE-1.7's 48 is the number that matters most here: shorter turn counts reduce both latency and the quadratic history-reprocessing cost that compounds silently across multi-turn sessions. The architecture — vertical RL applied to an open-weight Kimi K3 base — is now a reproducible template: license a strong open base, apply large-scale task-specific RL, pass efficiency gains through to users rather than capturing them as margin. The benchmark figures are Cognition's own; FrontierCode 1.1 Main results from independent labs are not yet available. What is observable immediately is that Cognition's Devin users get an 81% per-task cost reduction and faster convergence at the same price point — those are directly verifiable in production.

Verified across 1 sources: OfficeChaiAI

Agent Architectures & Tooling

Quadratic Chat History: 30-Turn Session Billed at 282,000 Tokens Despite 6,000 Typed — The Mechanism and the Four-Line Fix

A Friday technical post quantifies the hidden cost of stateless LLM APIs: in a 30-turn conversation with a 500-token system prompt, 200-token user messages, and 400-token replies, the input tokens billed total 282,000 — a 47x multiplier over the 6,000 tokens the user actually typed — because full conversation history is re-sent on every API call. Doubling conversation length from 30 to 60 turns increases costs nearly 4x. On 10,000 such conversations per month, the bill reaches $6,840 versus a naive estimate of $0.09 per conversation. A sliding-window fix — keeping the system prompt and last N exchanges, dropping earlier turns — reduces the 30-turn case to 143,400 input tokens (49% savings) with four lines of code; prompt caching adds a further layer but has its own break-even logic tied to cache TTL versus inter-request gap.

This failure mode is invisible in testing because short demo conversations look cheap and the quadratic curve only reveals itself at production scale. The concrete example — 6,000 typed tokens billed as 282,000 — is more operationally useful than generic cost-optimization advice because it gives a single diagnostic: log input-token count per request, sort descending, and the chat sessions sitting at the top of that list are the problem. The same dynamic applies to multi-step agent loops where accumulated tool outputs and prior reasoning are re-sent each turn. With DeepSeek's new off-peak rates, this cost compounds 100x cheaper per token but still quadratically with turn count — the fix is architectural, not a model swap.

Verified across 1 sources: Dev.to

Independent Print Publishing

USPS OIG: Five Years of 'Delivering for America' Cut Costs but Left Finances Unchanged — Rural Service Declined Most

Building on the $2.5 billion Q3 deficit and looming February 2027 cash shortage we've been tracking, a USPS Office of Inspector General audit released Tuesday found that the Delivering for America network transformation — a ten-year, $40 billion program launched in March 2021 — cut processing and delivery costs but failed to improve USPS's overall financial position. Rural communities experienced the sharpest service declines. The audit identified five systemic failures: no long-term cost-savings targets, no tracking of financial impact from network initiatives, absent program schedules, poor planning, and a fragmented facility network.

The core finding — that massive structural cost reduction did not translate into financial sustainability — means the pressure for further rate increases and the 90-cent stamp proposals we noted last month remain unmitigated. The explicit documentation of rural service decline compounds the Periodicals delivery issues we previously tracked in places like South Dakota. For Kav Magazine's P&L, the operational implication is that assuming predictable Periodicals delivery in upstate NY or Western Mass is not supportable by the agency's own metrics — distribution planning needs a margin for delay built in.

Verified across 2 sources: Center for Progressive Reform · CEP Research

Frum Community & Rockland Local

Agudath Israel and Becket Both File Supreme Court Amicus Briefs on Same Day Defending Home-Minyan Rights Under RLUIPA

Agudath Israel and the Becket Fund for Religious Liberty each filed amicus curiae briefs with the U.S. Supreme Court on Thursday in Grand v. City of University Heights, a case where Daniel Grand was told to obtain a special-use permit before hosting Shabbat minyans in his Ohio home. Agudath Israel argues that forcing religious Americans through discretionary permitting processes before courts can evaluate whether those processes chill core religious exercise violates RLUIPA's purpose; Becket makes the parallel procedural claim that Congress enacted RLUIPA specifically to prevent permit-exhaustion requirements from blocking judicial review. Oral arguments are scheduled for December 9, 2026. Grand withdrew his permit application and sued directly under RLUIPA after receiving a cease-and-desist letter from University Heights designating his home a 'place of religious assembly' under zoning code.

The question before the Court is not whether the permit requirement is substantively justified — it is whether Grand must exhaust the local permitting process before a federal court can hear his RLUIPA claim at all. A ruling that permit exhaustion is not required would give Orthodox communities immediate access to federal courts when zoning codes burden private religious gatherings, rather than spending years fighting city hall first. The direct implication for Monsey and Spring Valley is concrete: municipalities in Rockland County regularly apply residential-zoning conditions to home-based religious activity. A favorable ruling would shift the procedural posture of any such dispute, making federal RLUIPA litigation available at the outset rather than as a last resort.

Verified across 2 sources: Matzav · Jewish News Syndicate

Nyack Housing Authority Hearing Held on Voucher Expiration for Nine Orthodox Families — Section 8 Coverage Gap Documented at $4,300–$4,500 per Month

A Rockland County Supreme Court hearing was held Thursday on a lawsuit filed by nine Orthodox Hasidic families against the Nyack Housing Authority, alleging the authority allowed their Section 8 voucher eligibility to expire while they searched for four- and five-bedroom units. The families were restricted to Nyack Housing Authority's ZIP code 10960 service area, where two qualifying apartments available during their search rented at $4,300 and $4,500 monthly — above voucher coverage. Plaintiffs argue the authority failed to extend eligibility periods (available for up to 180 days under its administrative plan for large families) and failed to consider allowing county-wide searches. The families' correspondence invoked provisions prohibiting religious and familial-status discrimination; all nine are Orthodox Hasidic. Attorney Steven Yurowitz represents the plaintiffs; Keith Braunfotel represents the housing authority.

The documented market fact — two available large-family units priced at $4,300–$4,500, above what vouchers cover — quantifies a housing-supply gap that administrative flexibility alone cannot solve. Even if the court rules that the authority was required to extend eligibility periods and permit county-wide searches, the underlying supply problem remains: Rockland County lacks large-unit rental inventory within HUD Fair Market Rent limits. The case is now on the record as a disparate-impact claim for families with the specific size profile common in Orthodox communities, which sets a precedent for how housing authorities must administer the 180-day extension and geographic flexibility provisions. For a landlord with large units in Rockland County, this documents a structural demand pool — large Orthodox families with federal vouchers — that cannot currently be served within the existing HUD payment-standard framework.

Verified across 1 sources: Rockland News

Personal Finance Mechanics

8-Week Treasury Bill Prices at 3.845%, 20 bps Above Fed Ceiling — T-Bill Supply on Track to Double to $827B in 2026

Alongside the expanded $6B Treasury long-bond buybacks and yield pressures we've tracked over the past month, the U.S. sold eight-week Treasury bills on Thursday at a high rate of 3.845% — approximately 20 basis points above the Fed's 3.75% ceiling and 35 bps above the midpoint of the 3.50%–3.75% target range. This reflects both market pricing of a potential September FOMC hike and the Treasury's cash-account (TGA) buildup to approximately $950 billion. A Goldman Sachs projection estimates U.S. T-bill supply will double to approximately $827 billion in 2026, with bills now comprising 22% of total marketable debt, exceeding the Treasury's advisory committee's 15–20% comfort range.

Bill yields above the Fed's policy ceiling are a clean real-time read on market expectations, cleaner than futures pricing because bills mature within the policy decision cycle. The 3.845% rate on an 8-week bill means money-market funds and ultra-short ETFs tracking this end of the curve are now paying approximately 3.85% — competitive with high-yield savings accounts and directly relevant to how cash allocation decisions pencil against equity risk. The supply dynamic is the longer-term concern: $827 billion in annual bill issuance concentrated in weekly auctions creates a demand cliff at each rollover cycle. If money-market fund inflows slow or reverse — inflows were down $365 billion in H1 2026 per Ainvest — the Treasury either pays materially higher short rates or extends duration, both of which transmit directly into the rate environment facing small-landlord refinancing and acquisition financing.

Verified across 3 sources: aInvest · Ainvest · aInvest

Small Multi-Family Real Estate

Commercial Property Insurance Falls 6.3% in Q2 2026 — First Large-Account Decrease Since 2017 — While Umbrella Liability Rises for 35th Consecutive Quarter

WTW's Commercial Lines Insurance Pricing Survey (CLIPS) for Q2 2026 — covering 43 participating companies representing approximately 20% of the U.S. commercial insurance market — shows aggregate commercial insurance price increases of only 0.5%, down from 2.5% in Q1 2026 and 3.8% in Q2 2025. Commercial property saw the largest decrease, at 6.3%, with large commercial accounts experiencing their first price decrease since 2017. However, umbrella liability rose 5.3% for its 35th consecutive quarter and commercial auto rose 4.5%, driven by social inflation and litigation costs that outpace rate increases in those lines. The Economist independently corroborates the property-market softening and the liability-market divergence.

The bifurcation matters structurally: property coverage — the line where drone inspections, weather, and aging-building factors have driven multifamily operators' costs upward for years — is now in a soft market with new underwriting capacity entering, creating a genuine opportunity to rebid aggressively at renewal. Umbrella liability continuing to harden for 35 consecutive quarters is a separate problem that does not respond to market competition the same way, because nuclear verdict exposure and litigation-funding economics are the driver, not capital supply. Small multifamily landlords in Upstate NY and Western Mass should rebid property coverage now and review umbrella limits separately — the two lines are moving in opposite directions and require different strategies.

Verified across 2 sources: The Economist · InsurtechEye

Jewish History from the Archives

Sotheby's December Auction Features Joel ben Simeon Mahzor Estimated at $3–5M — Most Valuable Judaica Sale in House History

Sotheby's announced a December 17 Judaica auction headlined by a mahzor attributed to Joel ben Simeon — one of only 22 manuscripts attributed to him, with works held at the Library of Congress, Israel Museum, Jewish Theological Seminary, and British Library — dating to approximately 1475 Italy and estimated at $3–5 million. The sale also includes a 15th-century Italian Roman rite mahzor ($2–4 million), a monumental Hebrew Bible from circa 1300 Germany ($2–4 million), and an illuminated manuscript by Moses ben Jacob of Coucy ($1–2 million). Per Sotheby's senior Judaica specialist Sharon Liberman Mintz, the manuscripts 'trace the evolution of Hebrew manuscript production from medieval Ashkenaz to Renaissance Italy.' The works have been in private hands for approximately 100 years and will be exhibited at the National Library of Israel in November before the London auction.

Joel ben Simeon is among the very few medieval Jewish artists whose name survives attached to a documented body of work. The mahzor reportedly contains one of the earliest visual records of a Chanukah menorah and includes illustrations directly engaging with prayer text — evidence that medieval Jewish book culture was not purely scribal but involved active visual commentary on liturgy. The month-long National Library of Israel exhibition before auction is the practical access window for researchers; after December 17, these works enter private collections again. For anyone working on medieval Ashkenaz or Renaissance Italian Jewish material culture, this sale is a primary-source event, not a market report.

Verified across 1 sources: Jewish News Syndicate

Language & Etymology

Academy of the Hebrew Language: פָּרַד (Farad), Not Parak or Kalaf, Is the Correct Verb for Separating Pomegranate Seeds — Talmudic Corpus as Normative Source

The Academy of the Hebrew Language issued a lexicographic ruling this week correcting the common use of 'parak' (break apart) and 'kalaf' (peel) for separating pomegranate seeds, designating לְפָרֵד (le-farad) as the correct form. The Academy grounds its ruling in Mishnaic and Talmudic sources: פֶּרֶד or פְּרִידָה for the seed itself, and multiple attested constructions including פּוֹרֵט בְּרִימוֹן, הָרִימוֹן שֶׁפְּרָדוּ (the pomegranate whose seeds were separated), and מְפָרְדִין רִימּוֹנִים, demonstrating the root *parad* (separate) functioning in both Qal and Piel binyanim in the classical corpus.

The Academy's ruling is a case study in how Mishnaic Hebrew functions as a living normative authority for modern Israeli usage, with the standardization driven by multiple independent Talmudic attestations rather than intuition or popular practice. The root *parad* — shared with cognates meaning 'to separate' or 'to distinguish' across Semitic languages — is here given a domestic culinary application with explicit textual grounding. The timing, with Rosh Hashanah approaching and the pomegranate carrying its traditional 613-seed symbolism, makes this a lexicographic note with liturgical-cultural resonance. What the Academy's intervention does not provide is any usage-frequency data on how widely 'kalaf' and 'parak' are actually used versus 'farad' in spoken Israeli Hebrew — the correction may be philologically sound without being descriptively current.

Verified across 1 sources: KIPA

SMS & Low-Tech Product Design

Spain's SMS Sender-Registry Enforcement Is Live September 15 — 13,565 Registered, Thousands of Rejections Already

As the September 15 cutoff for Spain's SMS sender-registry approaches — which we tracked when enforcement was first announced — the CNMC registry had logged 13,565 inscribed aliases and approximately 5,000 in-process applications as of Thursday. The rollout has seen a massive volume of rejections due to missing owner authorization or failure to respond to regulator communications. Spanish telecommunications operators will begin automatically blocking SMS, MMS, and RCS messages using unregistered alphanumeric sender identities on Sunday, though the government is preparing a decree to allow in-process registrations to continue functioning during verification.

The high rejection rate — attributable to unopened regulator notifications and authorization-documentation gaps — reveals a compliance-awareness failure concentrated in microbusinesses without dedicated technical staff. Carrier-level automated blocking is a fundamentally different enforcement posture than financial penalties: there is no grace period and no appeal queue before the message fails to deliver. For any A2P SMS product targeting small businesses, this Spanish rollout joins the recent Vietnam and Kenya infrastructure enforcements in demonstrating that sender-identity verification is now a hard routing gate globally.

Verified across 2 sources: La Vanguardia · El Debate


The Big Picture

Cache Economics Are Diverging Faster Than Model Capability Gaps DeepSeek V4.1-Flash's $0.003/M off-peak cached-input rate arrives the same week that Anthropic's smart-reports feature and Fable 5.1 leaderboard results confirm Anthropic's competitive position is real — but only at 100–166x the cached-input cost. The BERI cost analysis and the quadratic chat-history piece both arrive at the same finding: headline benchmark scores are nearly irrelevant if the cache arithmetic doesn't pencil. The decision for agentic builders is no longer 'which frontier model' but 'which cache tier and at what input-frequency.'

Regulatory Gatekeeping Is Hardening Into Automatic Technical Blocks Spain's September 15 SMS sender-registry cutoff — where carriers will execute automated blocks on unregistered alphanumeric senders — and the USPS OIG audit's documentation of ongoing cost pressure that will produce further rate actions represent the same underlying pattern: regulators are moving from soft compliance guidance to infrastructure-level enforcement. Both hit small operators the hardest: the SMS blocks because microbusinesses never got notified, the USPS findings because periodicals publishers have no pricing power in response to rate increases.

Orthodox Community Legal Infrastructure Is Maturing Into Proactive Supreme Court Strategy Two independent amicus briefs filed on the same day — Agudath Israel's and Becket's — in Grand v. City of University Heights signal a coordinated legal campaign around RLUIPA's pre-enforcement review rights. Combined with the Section 8 voucher lawsuit against the Nyack Housing Authority, which implicates familial-size discrimination, and the SGO scholarship fund launching without New York or New Jersey opt-in, the Orthodox community is simultaneously pressing federal courts, federal housing policy, and federal tax architecture. These are not isolated cases — they share a common logic of establishing enforceable rights before local authorities can run out the clock.

Specialized Fine-Tuned Models on Open Weights Are Posting Frontier-Adjacent Scores at Fraction of Flagship Costs Cognition's SWE-2 — built on Kimi K3 and trained with a novel multi-effort RL method — scores 50% on FrontierCode 1.1 Main at 64% lower cost than Fable 5.1 and 81% fewer turns than its predecessor. This follows the same logic as HarnessDev (ByteDance) and PARSER (CUHK): the architecture and training regime, not the underlying parameter count, determine production performance. For practitioners, the implication is that open-weight bases plus vertical RL are now a credible alternative to closed flagship APIs for specific, well-scoped coding tasks.

Commercial Insurance Is Splitting Into Softening Property Lines and Hardening Liability Lines The WTW CLIPS data and the Economist's corroborating analysis show property insurance prices falling 6.3% in Q2 2026 — genuine relief for multifamily operators — while umbrella liability rose 5.3% for its 35th consecutive quarter and commercial auto continued upward. The Florida condo insurance situation, where structural reserve requirements now trigger policy withdrawals on aging stock, points at a regulatory model that could migrate to other states. Small landlords in Upstate NY and Western Mass need to rebid property coverage aggressively now while the market is soft, and review umbrella adequacy separately.

What to Expect

2026-09-14 DeepSeek automatically routes all V4-Pro API requests to V4.1-Flash at the lower price point — existing production deployments need regression testing before this date.
2026-09-15 Spain's CNMC enforces alphanumeric SMS sender-ID registration cutoff: carriers begin blocking unregistered business senders. Also: PRC public comment deadline for USPS Negotiated Service Agreement dockets MC2026-374 through MC2026-378 (Priority Mail Express/Priority Mail/Ground Advantage contract 1511).
2026-09-15 FOMC meeting begins (September 15–16); 17-week bill pricing at 3.895% — 14 bps above the Fed's 3.75% ceiling — is the market's current read on hike probability.
2026-09-17 New York State Public Health Council votes on the $2.245 billion Maimonides Medical Center transfer to NYC Health + Hospitals — four Hasidic congregations and a 34,000-signature petition are on record in opposition.
2026-12-09 Supreme Court oral arguments in Grand v. City of University Heights (RLUIPA home-minyan permitting case); Agudath Israel and Becket both filed amicus briefs September 10.

Every story, researched.

Every story verified across multiple sources before publication.

🔍

Scanned

Across multiple search engines and news databases

909
📖

Read in full

Every article opened, read, and evaluated

185

Published today

Ranked by importance and verified across sources

12

— The Primary Source

🎙 Listen as a podcast

Subscribe in your favorite podcast app to get each new briefing delivered automatically as audio.

Apple Podcasts
Library tab → ••• menu → Follow a Show by URL → paste
Overcast
+ button → Add URL → paste
Pocket Casts
Search bar → paste URL
Castro, AntennaPod, Podcast Addict, Castbox, Podverse, Fountain
Look for Add by URL or paste into search

Spotify isn’t supported yet — it only lists shows from its own directory. Let us know if you need it there.