Today on The Bandwidth-Bound: Engineers are systematically hardening low-level execution pathways as sub-quadratic architectures scale up. Across Apple Silicon and CUDA, maintainers are actively dismantling dispatch overheads and KV-cache alignment penalties to protect decoding throughput, while orchestration frameworks increasingly rely on direct environmental execution to verify multi-agent loops.
A GitHub issue report filed on Monday, October 5, 2026, details why enabling prefix caching (`mamba_cache_mode="align"`) for hybrid Mamba2 models incurs an 11% to 16% throughput penalty on NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 across 4x GB200 nodes. Profiling reveals that the slowdown is caused by per-group eager block-table operations, per-step alignment kernels running outside CUDA graphs, and prefill splits at block boundaries that duplicate fixed non-GPU overhead. Engineers are actively drafting PRs to fuse alignment kernels and share batch metadata across KV cache groups.
Why it matters
Prefix caching is essential for multi-turn agent serving and long-context prompt reuse, but naive alignment modes in hybrid SSM architectures reintroduce severe host-bound CPU overheads. For local-LLM practitioners and systems developers, these profiling breakdowns demonstrate that architectural memory savings can be undermined by unoptimized CUDA launch paths. Resolving these kernel-graph boundaries is necessary to achieve theoretical linear-attention speedups under continuous batching.
Inference engineers on the issue note that executing alignment steps outside CUDA graphs introduces catastrophic launch serialization on Blackwell systems. Maintainers argue that while sharing metadata across KV groups reduces host overhead, maintaining strict state alignment across hybrid recurrent layers requires specialized fused kernels rather than standard vLLM block-table abstractions.
On Monday, October 5, 2026, an issue filed against the Saragossa Metal inference engine analyzed why running Qwen3.8-27B dense u4 achieves ~30 tok/s during short decoding compared to 40-80 tok/s in `mlx-lm`. Dispatch tracing showed roughly 450 GPU dispatches per token—averaging 7 separate dispatches per layer outside attention and DeltaNet kernels—which serializes Metal command queues and starves memory bandwidth. Proposed optimizations focus on porting u8 fusion patterns, combining RMS norms with residual additions, and fusing DeltaNet convolution, gating, and scan loops into single kernels.
Why it matters
Executing hybrid models efficiently on Apple Silicon requires aggressive operator fusion, especially when stepping between dense MLPs, full attention, and recurrent DeltaNet layers. Unmerged operations convert memory-bandwidth-bound decoding steps into latency-bound CPU-launch queues, significantly reducing generation throughput. For local-LLM developers, these profiling metrics highlight why custom Metal backends must fuse normalizations and gating steps to match established frameworks.
Saragossa engine contributors emphasize that dispatch serialization is the primary wall preventing full utilization of Apple Silicon memory bandwidth on hybrid quants. Conversely, `mlx-lm` maintainers maintain that higher-level graph execution automatically coalesces these passes, though low-level custom C/Metal runtimes require explicit hand-written kernel fusion to eliminate command-buffer latency.
On Monday, October 5, 2026, Anthropic's interpretability research team published findings showing that Claude spontaneously organizes internal representations into a global workspace structure. Using a multivariable calculus diagnostic tool called the Jacobian Lens (J-lens), researchers identified a low-dimensional computational control bus ('J-space') that acts as a working memory buffer across distributed attention heads. The study confirms that standard cross-entropy pre-training drives transformer layers to centralize state routing through this low-rank subspace.
Why it matters
Identifying a low-dimensional control bus inside autoregressive transformers provides mathematical evidence of convergent architectural structure in deep models. For mechanistic interpretability researchers, J-space provides a concrete target for activation patching, probing, and real-time state monitoring. Mapping this control register allows researchers to inspect model working memory directly rather than relying on high-dimensional residual stream approximations.
Anthropic researchers frame J-space as a functional analog to biological Global Workspace Theory, proving that predictive objective functions naturally discover centralized information routing. Independent interpretability researchers caution that while Jacobian Lens readouts isolate low-rank control vectors, causal intervention studies are still required to prove J-space steering does not induce collateral performance degradation.
Following our coverage of the Claude Code unsandboxed TypeScript 'Mods' framework rollout in version 2.1.287, Anthropic pushed v2.1.289 on Saturday, October 3, 2026. This rapid follow-up introduces rendering and asynchronous exception firewalls for the new extension layer, preventing single-extension crashes from terminating main CLI sessions. The release also adds `agent.spawn` to the plugin API, allowing extensions to initiate multi-agent parallel workflows, and patches permission bypass vulnerabilities related to environment variable prefix overrides in sandboxed commands.
Why it matters
As Claude Code transitions into an extensible middleware platform—a shift we've tracked across recent releases—unhandled extension errors previously represented a single point of session failure. Adding fault-isolation firewalls and programmatic sub-agent spawning stabilizes the long-running autonomous execution loops now moving into production. For developers building agent tooling, `agent.spawn` provides a clean API for triggering parallel review and verification sub-agents directly within custom mods.
Anthropic engineers highlight that isolation firewalls protect terminal UI state during complex background tool executions. Security researchers note that while environment variable overrides were patched, running unsandboxed TypeScript mods still requires rigorous local code auditing to prevent credential leakage.
GitHub issue #410 in the `etalii.dllm` repository detailed the Phase 69 release on Sunday, October 4, 2026, introducing Sparse Autoencoders (SAEs) and editing tooling for T5 text-to-text architectures. The update includes interpretability documentation, getting-started guides, golden fingerprints for T5 activation edits, and gradient validation tests comparing encoder residual gradients against finite differences and PyTorch autograd. The feature suite was written using Claude Code.
Why it matters
Expanding mechanistic interpretability tooling to encoder-decoder architectures like T5 requires verified gradient checks and reproducible baseline runs. By publishing golden fingerprints and gradient validation against finite differences, the project provides open-source assets for researchers extending personal probing toolkits. This release simplifies feature extraction and activation editing on non-causal language models.
The maintainers note that providing exact golden values ensures researchers can verify local SAE probe installations without numerical drift. Independent interpretability practitioners appreciate the inclusion of finite-difference gradient checks to validate custom hook hooks across PyTorch versions.
A paper published Thursday, October 1, 2026, by researchers at Cambridge and Google Cloud AI introduced VeriHarness, a verification framework for long-horizon autonomous tasks. The study demonstrates that majority voting and self-consistency consensus across multiple rollouts frequently hide correlated errors, whereas rollout disagreement reliably highlights minority paths containing correct answers. By pairing an adversarial dispute resolver with an execution environment and incremental token caching, VeriHarness achieved average score gains of 6.2 points on Gemini 3.5 Flash and 6.4 points on Claude Opus 4.8 across five long-horizon benchmarks.
Why it matters
Relying on majority-vote consensus or LLM-as-a-judge scoring creates a false sense of security in agent orchestration, as models systematically share common failure modes. VeriHarness shifts the verification paradigm toward execution-grounded environment checks and dispute resolution. For developers building multi-agent verification loops, this approach offers a clear method for extracting correct code diffs without relying on external reference answers.
The paper's authors argue that agreement among generative models is a flawed proxy for truth in long-horizon reasoning, making active environmental execution mandatory. External framework engineers note that while consensus-challenging verifiers improve accuracy, managing parallel rollout execution and incremental token caches adds non-trivial latency and compute costs.
A paper published Monday, October 5, 2026, by Microsoft researchers introduced Agensh, a self-organized, multi-agent coding harness that eliminates central orchestrator nodes. Sub-agents asynchronously claim tasks, run local tests, merge code into a shared Git repository, and log state updates on an open board. Evaluated on ProgramBench using GPT-5.6-sol high, scaling from 1 to 128 agents raised mean test-pass rates from 19.31% to 28.78%, while scaling to 1,024 agents on the `pandoc` benchmark increased pass rates from 33.89% to 55.06%.
Why it matters
Centralized orchestrator models encounter context window limits and sequential bottlenecks when managing large codebases. Agensh proves that fully decentralized, peer-to-peer agent fleets can coordinate via standard Git repositories and shared state boards to solve complex engineering tasks. This architecture offers a blueprint for scaling multi-agent developer workflows without relying on a single, expensive lead-agent router.
The authors emphasize that removing central controllers bypasses context bloat and enables linear task scaling across massive repositories. Independent agent developers note that while test-pass rates increase significantly at 1,024 agents, the token expenditure and API costs scale rapidly, making execution-based verification essential to prevent merge conflicts.
On Monday, October 5, 2026, a technical proposal submitted to the PrimeIntellect PrimeRL repository introduced lossless sparse delta synchronization for distributed reinforcement learning training loops. Citing empirical measurements showing that only ~1% of weight elements change between consecutive RL PPO steps, the framework transfers compact index-encoded weight deltas over WAN connections instead of full model parameters. The architecture features multi-stream pipelining, two-tier fan-out relays, and automated background recovery for rollout workers.
Why it matters
Distributed post-training and RL alignment for open-weight models are frequently bottlenecked by multi-gigabyte weight broadcasts between trainer nodes and vLLM rollout workers across cloud regions. Syncing only sparse delta indices reduces network transfer overhead by over 90%, preventing GPU idling during iterative alignment runs. This network optimization accelerates local and multi-cloud agent training cycles.
PrimeIntellect engineers demonstrate that sparse delta encoding eliminates network saturation during high-frequency PPO steps. Distributed systems researchers point out that while delta synchronization cuts WAN bandwidth, rollout nodes must maintain atomic state verification to prevent drift if a delta packet drops.
On Monday, October 5, 2026, researchers introduced VERSE (Verified Self-Evolving optimizer), a system that automatically modifies agent prompts, tools, and execution hooks while keeping underlying model weights frozen. VERSE uses execution-based environment verification to test draft harness edits, replay failed runs, and perturb step sequences. Tested on held-out SWE-rebench tasks across five programming languages, VERSE's validation-selected harness achieved 42.3% and 37.7% accuracy compared to 39.2% and 29.3% for unverified baseline optimizers.
Why it matters
Automated prompt and harness optimization often degenerates into benchmark memorization when guided purely by text-based feedback. VERSE demonstrates that execution-grounded verification gates allow optimizers to safely construct custom tools and revise execution flows without touching weight parameters. This provides a systematic framework for tuning open-weight agent scaffolds for domain-specific coding environments.
The paper's authors highlight that execution feedback is mandatory to prevent self-evolving optimizers from generating brittle prompt hacks. Framework developers note that while VERSE improves out-of-distribution task performance, running iterative execution verification loops increases offline harness training budgets.
On Sunday, October 4, 2026, developers released Agenthof v0.1.0 (Apache-2.0), a control plane designed to govern autonomous LLM agents by routing all model actions through explicit execution 'doors' (`model`, `tool`, `exec`, and `spawn`). The system authorizes API calls, injects credentials that the model never directly reads, and logs tamper-evident append-only ledger events. Built as a framework-neutral layer supporting HTTP and MCP endpoints with a light LangChain wrapper, unauthorized model attempts result in logged refusals.
Why it matters
Granting autonomous agents direct access to raw environment credentials introduces critical security vulnerabilities regarding prompt injection and secret exfiltration. Agenthof isolates credentials within an outer control plane, injecting tokens only at execution boundaries while recording tamper-evident audit chains. This architecture enables secure local and cloud execution of unconstrained tool-calling agents.
The framework authors argue that agent security must be enforced by outer infrastructure doors rather than trusting model system prompts. Security auditors favor append-only ledger tracking, though framework developers observe that proxying all MCP tool calls through credential doors introduces minor latency overhead.
A detailed bug report (Issue #16000) filed against the `t3code` repository on Monday, October 5, 2026, revealed that altering a thread's runtime mode, worktree directory, or model provider instance abruptly detaches active Claude sessions. The transition handler terminates background sub-agents and shell processes without checking for pending task completion or active locks, bypassing session safety guards and causing unrecoverable state loss.
Why it matters
Orchestration harnesses managing asynchronous sub-agents require robust process lifecycle state machines during configuration mutations. Terminating background shell workers during thread switching leads to silent code corruption and dangling git locks. For agent tool developers, this issue highlights the necessity of implementing process drain checks prior to updating runtime context.
The bug reporter showed that context switches immediately kill active sub-agent PIDs without emitting clean exit signals. Harness maintainers acknowledged the flaw, proposing that state mutations should be queued behind an asynchronous sub-agent completion barrier.
Developer StayLameBro released Backburner v0.0.1 on Friday, October 2, 2026, an open-source tool connecting an iPhone or iPad to an Apple Silicon Mac via a 10 Gb/s USB-C cable to distribute inference compute. Utilizing a modified `llama.cpp` fork, the system offloads long prompt prefill tasks and secondary KV-cache storage to the iOS device. Benchmarks running Qwen3.8-27B show prompt prefill speedups of 29% to 44% while extending effective context capacity from 64k to 140k tokens under 4-bit quantization.
Why it matters
Local execution of 27B+ parameter models on entry-level Macs is heavily constrained by unified memory capacity and bandwidth limits. Backburner introduces a practical heterogeneous execution model, repurposing idle mobile chips into secondary matrix-multiplication units and KV-cache offload banks. This provides local-LLM practitioners with a zero-cost hardware extension path for running long-context models without purchasing dedicated workstation upgrades.
The developer demonstrates that high-speed USB-C interfaces provide sufficient interconnect bandwidth to make distributed tensor execution viable between Apple devices. Community testers observe that while prefill and context limits improve substantially, decode throughput remains bounded by the primary Mac's memory bandwidth.
Verified across 2 sources:
Traictory(Oct 4) · KOCPC(Oct 5)
Click Copy for AI above, then paste the prompt
into your favorite AI chatbot — ChatGPT, Claude, Gemini, or
Perplexity all work well.
A GitHub tracking issue published Monday, October 5, 2026, for the `vllm-mlx` fork detailed an active research round targeting execution bottlenecks on Apple Silicon. Key investigation areas include probing wired-memory residency decay during GPU idle pauses, evaluating Metal command-buffer constraints under sparse MoE routing, and testing dequantize-then-matmul passes for dense prefill chunks. The update also audits shortest-prefill-first scheduling and memory-pinning patches for MLX versions 0.32.3 and 0.32.4.
Why it matters
Running open-weight MoE models on Apple Silicon hardware often suffers from unpredicted latency spikes caused by Metal command-buffer limits and OS-level memory unpinning during idle steps. Systematically addressing memory residency decay and prefill chunking directly stabilizes generation performance. Tracking these engine patches helps local practitioners eliminate cold-start penalties in desktop agent runtimes.
Fork maintainers argue that bypassing Metal command-buffer serialization is critical for maintaining steady token throughput on Mac Studio clusters. Systems developers note that while shortest-prefill-first scheduling improves interactive responsiveness, heavy concurrent prefill jobs can still starve background decoding steps.
A GitHub issue filed in the `llama.cpp` repository on Monday, October 5, 2026, reported a 50% performance drop in prompt evaluation speeds for Qwen3.6-35B-A3B-GGUF when running `llama-server`. Bisecting the regression traced the cause to commit `bed0a85` (merged via PR #29184), which fused shared experts into MMVQ kernels for CUDA backends. While standard `llama-bench` runs did not reveal the bug, end-to-end completion tests showed prompt evaluation throughput dropping from ~13,816 tok/s to ~6,858 tok/s on a system running triple NVIDIA RTX PRO 6000 Blackwell GPUs.
Why it matters
Low-level CUDA kernel fusions designed to accelerate matrix-vector operations can introduce severe performance regressions during batched prompt prefill steps. This issue underscores why synthetic micro-benchmarks like `llama-bench` are insufficient for catching real-world serving degradation. For local-LLM operators, identifying this commit prevents silent prefill bottlenecks in production MoE serving pipelines.
The issue author demonstrated that synthetic benchmarks failed to catch the regression because they do not simulate multi-query completion server contexts. Kernel maintainers acknowledged that while MMVQ fusion benefits single-token decode passes, it introduces kernel launch bottlenecks during large prompt ingestion batches.
On Monday, October 5, 2026, Salvatore Sanfilippo detailed performance metrics for DS4 (DwarfStar 4), a lightweight C inference engine designed for sparse MoE models. Using an asymmetric 2/8-bit quantization scheme, DS4 fits the 284B-parameter DeepSeek V4 Flash model into ~96.5GB of memory. Benchmarks demonstrate decode speeds reaching 39.4 tok/s on a 128GB MacBook Pro M5 Max and 32 tok/s on AMD Strix Halo hardware using custom forks.
Why it matters
Executing 200B+ Mixture-of-Experts models locally typically requires multi-node GPU clusters due to strict VRAM footprint constraints. DS4 proves that asymmetric quantization—compressing non-critical expert layers down to 2 bits while retaining 8-bit precision on attention projections—allows near-frontier MoE architectures to run on single unified-memory workstations. This lowers the hardware bar for offline local inference.
Sanfilippo highlights that custom C runtimes stripped of framework abstraction achieve significantly higher token throughput on Mac hardware. Community developers note that while asymmetric quants preserve general response quality, extreme 2-bit expert compression requires testing against domain-specific code benchmarks.
A pull request proposal submitted on Sunday, October 4, 2026, introduces an opt-in `--cache-type-kv int8` parameter to the GGUF server codebase, specifically tailored for Qwen3.8 Flash-Next. The implementation applies per-head and per-block INT8 scaling to full-attention KV matrices, cutting memory scaling factors in half and doubling concurrent 256K context sessions. Fused quantization and dequantization kernels were written to avoid memory bandwidth penalties on hardware lacking native FP8 instructions, with verification gates requiring logit MSE checks and Needle-In-A-Haystack validation up to 256K context depths.
Why it matters
At extreme context lengths like 256K, the memory required for KV-cache state quickly surpasses the base model weight footprint. Scoping INT8 KV quantization directly to Qwen3.8's hybrid architecture allows operators to double session concurrency without relying on uncalibrated global precision drops. For local inference engineers, fused quantization kernels ensure that cache compression improves VRAM capacity without degrading decoding speeds on consumer GPUs.
The PR authors emphasize that per-block scaling is necessary to prevent severe accuracy loss in deep attention heads during long-context retrieval. Engine maintainers highlight that requiring explicit quality gates—such as logit MSE and needle-in-a-haystack passes—prevents broken quantization kernels from being merged into release builds.
A paper published Sunday, October 4, 2026, presented HeteroFold, a method for transferring key-value (KV) caches directly between different model families without re-running prefill steps or updating model weights. The process aligns architectural layer structures, projects the source cache into the target model's latent space, and applies behavior calibration. Experiments showed that transferring a 32K context cache from Llama-3.1-8B to Ministral-3-14B ran 10.7 times faster than native target prefill, matching text-based context passing accuracy on multi-agent benchmarks.
Why it matters
Passing long context across heterogeneous multi-agent pipelines currently forces every receiving model to waste compute re-processing prompt prefill. HeteroFold eliminates redundant prefill compute by translating KV-cache representations directly between model architectures. For multi-agent systems, this cross-model cache transfer drastically reduces turn latency during high-frequency message passing.
The authors show that latent space projection preserves long-context retrieval capabilities across model boundaries without weight retraining. Inference engineers note that while 10x prefill speedups are impressive, managing cross-model alignment matrices adds complexity to serving engine KV-cache managers.
A paper published Monday, October 5, 2026, presented Hyperbolic Entropy Steering (HEST), a framework that embeds language model hidden states into the Poincaré ball using a lightweight probe guided by next-token entropy. When token generation entropy crosses a defined threshold, HEST shifts the activation state along the geodesic of steepest descent before mapping it back to the Euclidean residual stream. Tested on Qwen2.5-Math and Llama-3.1 checkpoints, HEST improved greedy accuracy on MATH-500 and GSM8K by up to 1.8 points without degrading output quality on low-entropy tokens.
Why it matters
Standard Euclidean activation steering applies constant steering vectors across all tokens, which frequently corrupts generation during straightforward sequences. By leveraging hyperbolic geometry to match the natural tree structure of reasoning paths, HEST applies interventions strictly when the model exhibits hesitation. This provides interpretability researchers with a targeted, non-Euclidean methodology for dynamic runtime model steering.
The authors show that probing hidden states in non-Euclidean space captures branching decision points far better than linear probes. Mechanistic interpretability practitioners note that while entropy-triggered interventions prevent over-steering, calibrating probe thresholds across diverse fine-tunes requires automated per-model tuning.
Verified across 2 sources:
Prismix(Oct 5) · arXiv(Oct 5)
Click Copy for AI above, then paste the prompt
into your favorite AI chatbot — ChatGPT, Claude, Gemini, or
Perplexity all work well.
On Monday, October 5, 2026, developer Niko1221 detailed updates to the MIT-licensed Strata inference engine, demonstrating execution of Qwen3.8-Flash-Next (125B total parameters, Q2_0 quantization) on a system with a 12GB VRAM RTX 5070 GPU, 32GB RAM, and an NVMe SSD. Strata uses a three-tier memory scheduler that holds active expert layers in VRAM, manages misses in system RAM, and reads an n-gram table off storage, achieving 94 tokens per second during decode steps by combining custom CUDA/HIP kernels with speculative decoding.
Why it matters
Memory capacity walls have traditionally prevented local execution of 100B+ Mixture-of-Experts architectures without multi-GPU servers. Strata's three-tier scheduling pipeline proves that VRAM, system RAM, and NVMe storage can be unified into an efficient inference hierarchy. For local-LLM practitioners, this framework enables offline execution of frontier-scale sparse models on standard desktop hardware.
The developer highlights that speculative decoding and n-gram lookahead tables effectively mask NVMe and RAM transfer latencies during autoregressive steps. Systems researchers point out that while decode speeds hit 94 tok/s under heavy quantization, prefill phases remain severely bottlenecked by PCIe and system RAM bandwidth.
On Monday, October 5, 2026, Jared Palmer released Kev 1.0, a suite of open-weight decision models fine-tuned from Qwen3.5 and Qwen3.8 base architectures. Ranging from 0.8B parameters (runnable on 4GB VRAM) to 27B parameters, the models are trained to handle classification, multiple-choice, and rating tasks in a single forward pass with calibrated output probabilities. The release includes full checkpoints on Hugging Face, local MLX serving scripts for Apple Silicon, and an automated fine-tuning skill built for Modal.
Why it matters
Routing agent workflows through multi-gigabyte general-purpose LLMs introduces unnecessary latency and expense for basic decision points. Kev provides compact, open-weight classification models calibrated specifically for structured JSON decision outputs. For local practitioner pipelines, running 0.8B to 4B decision models natively via MLX or CUDA significantly cuts system latency during agent tool routing.
The author notes that calibrating output probabilities directly in small base models eliminates the need for expensive multi-token reasoning chains on binary routing choices. Framework developers praise the inclusion of native MLX serving scripts, though some note that domain-specific classification tasks still require task-specific fine-tuning.
Kernel Dispatch and Alignment Overheads Cap Hybrid Linear Attention Throughput While Gated DeltaNet and Mamba2 layers reduce theoretical KV-cache footprints, real-world serving engines are encountering host-bound execution limits. Unfused dispatch chains on Metal and per-group eager block operations in CUDA graphs introduce latency spikes that cancel out mathematical memory savings during decoding steps.
Execution-Grounded Adjudication Replaces Self-Reported Agent Completion Multi-agent frameworks are abandoning majority voting and self-certified completion in favor of isolated verifier agents and tool-backed execution checks. Empirical data shows model consensus masks correlated errors, pushing developers to adopt environment execution and adversarial dispute resolution before accepting code diffs.
Heterogeneous Edge Hardware Pairs Mobile Devices for Context Expansion To bypass strict unified memory and VRAM walls without purchasing enterprise GPUs, open-source runtimes are offloading context prefill and KV-cache storage across high-speed USB-C interfaces to secondary mobile chips. Offloading non-active state to external devices extends local context limits past 140k tokens.
Dynamic and Scoped KV-Cache Quantization Hardens Long-Context Serving Engine developers are moving away from blanket weight-quantization rules toward model-specific, per-head INT8 and MXFP4 KV-cache schemes. By pairing targeted cache compression with fused kernels, serving runtimes double concurrent long-context sessions without triggering generation divergence.
Probing Tools Target Entropy Junctions for Non-Euclidean Steering Mechanistic interpretability is moving toward non-Euclidean state steering, embedding LLM residual states into hyperbolic space to intervene specifically at high-entropy decision branches. This targeted activation editing prevents global model degradation while improving accuracy at hesitation points.
What to Expect
2026-10-15—Expected upstream vLLM MRv2 merge window for unified Metal KV storage adapters.
How We Built This Briefing
Every story, researched.
Every story verified across multiple sources before publication.
🔍
Scanned
Across multiple search engines and news databases
454
📖
Read in full
Every article opened, read, and evaluated
116
⭐
Published today
Ranked by importance and verified across sources
20
— The Bandwidth-Bound
🎙 Listen as a podcast
Subscribe in your favorite podcast app to get each new briefing delivered automatically as audio.
Apple Podcasts
Library tab → ••• menu → Follow a Show by URL → paste