🧪 The Bandwidth-Bound

Thursday, October 1, 2026

19 stories · Deep format

Generated with AI from public sources. Verify before relying on for decisions.

🎧 Listen to this briefing or subscribe as a podcast →

We've spent the past few weeks tracking the rapid adoption of hybrid linear attention architectures. Now, open-source kernel maintainers and serving frameworks are zeroing in on the architecture's next major hurdle: recurrent memory. Rather than relying purely on hardware scaling, developers are deploying windowed state quantization and dynamic precision allocation to stabilize long-context performance.

Linear & Hybrid Attention Architectures

LeapQuant Delivers Lossless 8-Bit Recurrent-State Quantization for Linear Attention

On Tuesday, September 29, 2026, researchers from UC Berkeley, UW, MIT, Perplexity AI, and NVIDIA published LeapQuant, a training-free 8-bit recurrent-state quantization scheme for linear attention models. The method applies per-window quantization to prevent error accumulation across token steps and inserts high-precision Compensator Tokens for state outliers. Across Qwen, Kimi, and GLM hybrid architectures, LeapQuant yielded kernel speedups of 2.05–3.70x and an end-to-end decode speedup of 1.47x on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs while cutting state memory traffic by 3.4x.

While sub-quadratic hybrid architectures bypass the quadratic growth of standard attention KV caches, their fixed-size recurrent states incur heavy high-bandwidth memory (HBM) read/write traffic during decoding. LeapQuant resolves the numerical collapse that usually occurs when quantizing these state matrices without requiring costly fine-tuning or architectural modifications. For practitioners deploying long-context open-weight hybrids, this provides a drop-in mechanism to lower HBM memory bandwidth pressure.

The authors emphasize that window-based quantization combined with outlier tokens successfully stabilizes autoregressive generation across 100k+ sequence lengths. However, systems engineers note that implementing custom compensator token kernels requires specialized GPU dispatch paths that are not yet natively integrated into default Triton or vLLM runtimes.

Verified across 1 sources: arXiv (Sep 29)

STEPQuant Applies Time-Space Precision Allocation to Recurrent State Models

A paper published Tuesday, September 29, 2026, introduced STEPQuant, a post-training quantization framework for delta-rule recurrent states in linear attention architectures. The method analyzes the error magnitude and lifespan of recurrent memory components, jointly fitting key-row and value-column scales to assign dynamic precision. Applied to Qwen and Kimi linear models inside SGLang, STEPQuant achieved FP32-level accuracy at 6-bit precision and outperformed uniform INT8 quantization at 4-bit, reducing overall serving memory usage by up to 68.7%.

Uniform low-bit quantization fails on linear recurrent states because state errors compound forward across every generated token. By identifying that long-lived memory channels carry disproportionate sensitivity, STEPQuant allows serving engines to allocate 8-bit or FP16 budgets selectively while squeezing less sensitive states down to 4-bit. This optimizes the memory footprint for long-sequence serving on hardware with restricted VRAM capacity.

The research team highlights that joint key-value scaling preserves sequence perplexity across multi-turn chats where standard PTQ methods degrade. Conversely, independent open-source maintainers caution that dynamic precision allocation introduces state layout fragmentation inside custom CUDA attention kernels.

Verified across 3 sources: GitHub (Oct 1) · arXiv (Sep 29) · Crunch AI Research (Sep 30)

SGLang v0.5.19 Ships Deployment Configurations and Memory Ratio Calculations for Qwen3.8-27B

On Thursday, October 1, 2026, SGLang documentation was updated for version 0.5.19, detailing exact serving configurations for Qwen3.8-27B dense hybrid Gated Delta Networks (GDN). The release includes explicit formulas for the `--mamba-full-memory-ratio` parameter, which splits post-weight GPU VRAM between the GDN recurrent state pool and the paged attention KV cache pool across its 48 linear layers and 16 full-attention layers. Tested configurations validate GSM8K scoring across RTX 5090, RTX PRO 6000, and DGX Spark nodes under NVFP4 and FP8 execution.

Deploying hybrid architectures that combine 3:1 linear-to-full attention ratios requires explicit memory partitioning to prevent runtime out-of-memory errors under high request concurrency. Providing hardware-validated sizing formulas and Triton prefill flags allows local practitioners to maximize request batching without starving either the paged KV cache or the recurrent state allocations.

SGLang maintainers assert that explicit ratio controls eliminate static memory waste and support higher concurrency on unified-memory platforms like DGX Spark. Infrastructure operators note, however, that tuning `--mamba-full-memory-ratio` requires manual profiling per prompt-length distribution.

Verified across 1 sources: SGLang Documentation (Oct 1)

Proposal Recommends Chunkwise Recurrence and Tiled Prefill for Qwen 3.5 Hybrid Gated DeltaNet

A detailed technical proposal posted to GitHub on Wednesday, September 30, 2026, outlined three performance enhancements for Qwen 3.5 hybrid Gated DeltaNet models, validated via a standalone ROCm engine (`rocml`). The proposal presents a blocked chunkwise Gated Delta Net prefill path that cuts GDN recurrence prefill overhead from 30% to 13%, a tiled online-softmax prefill kernel, and a KIVI-style sink-plus-window quantized bulk KV cache with inline dequantization.

Sequential state updates in Gated DeltaNet layers create a serial bottleneck during the prefill phase of long prompts. Chunking the linear recurrence computation into parallel blocks and combining it with inline KV-cache dequantization allows hybrid architectures to maintain high prefill throughput without swamping GPU memory bandwidth.

The author demonstrates that chunkwise prefill transforms sequential state updates into parallel matrix multiplications. ROCm maintainers note that while the chunkwise path delivers substantial speedups on AMD hardware, porting the tiled online-softmax kernels to Triton requires separate register tuning for NVIDIA GPUs.

Verified across 2 sources: GitHub (Sep 30) · GitHub (Sep 30)

RFC Proposes Native Rust and CUDA Inference for Mapika Decider-2B Hybrid Model

An RFC issue filed on GitHub on Thursday, October 1, 2026, proposed adding native Rust and CUDA inference support for Mapika/decider-2b, a 2B parameter decision model utilizing 18 Gated DeltaNet layers and 6 full-attention layers. The implementation bypasses standard generative token generation to score allowed prompt choices directly across choice, score, and noul templates, re-using Qwen 3.5 CUDA infrastructure inside the System1-Omni engine.

Specialized decision models built on hybrid linear-attention architectures perform best when executed via direct log-probability scoring paths rather than full text generation loops. Building native Rust and CUDA scoring workers removes Python execution overhead and enables low-latency decision routing for agent frameworks.

The proposal author notes that direct logit scoring on allowed choices cuts execution latency by over 80% compared to autoregressive generation. Framework maintainers stress that choices must be strictly bounded to prevent index out-of-bounds errors during direct logit extraction.

Verified across 1 sources: GitHub (Oct 1)

Open-Weight Model Releases

IQuest Releases IQuest-Q1 320B Sparse MoE with Native 512K Context

On Wednesday, September 30, 2026, IQuest open-sourced IQuest-Q1, a 320-billion total parameter Mixture-of-Experts model that activates roughly 15 billion parameters per token via a 256-expert sparse routing scheme (8 active per forward pass). The model features 88 transformer layers configured in a 3:1 sliding-window to full-attention ratio and partial RoPE across a native 512K context window. It ships with native multi-token prediction (MTP) modules for speculative decoding and direct adapter integrations for Claude Code and Codex CLI.

IQuest-Q1 demonstrates how sparse MoE routing can deliver high capacity while maintaining manageable active parameter counts for local or single-node inference. The inclusion of native MTP modules and pre-built harness configs for tools like Claude Code highlights a shift toward shipping open-weight models ready for immediate deployment in developer agent workflows.

IQuest developers emphasize that the 256-expert fine-grained routing minimizes parameter redundancy while retaining reasoning depth. Independent benchmarks express concern that managing 256 separate expert parameter weights creates significant NVMe and VRAM offloading overhead on non-enterprise GPU setups.

Verified across 1 sources: MindStudio (Sep 30)

NaiveAI Drops Naive-N0.5-Flash 309B MoE Replacing Full Attention with Hybrid SWA-DSA

On Wednesday, September 30, 2026, NaiveAI released Naive-N0.5-Flash under an MIT license, an open-weight MoE model built on Xiaomi's MiMo base featuring 309B total and 15.5B active parameters per token. The architecture completely eliminates standard quadratic full-attention layers, relying instead on 48 transformer layers alternating between Sliding-Window Attention (SWA) and DeepSeek Sparse Attention (DSA) to achieve a native 1-million-token context window. The release includes NaiveRT, a custom inference backend optimized for long-context execution.

Completely removing full-attention layers in favor of sparse indexing mechanisms like DSA provides an architectural route to scale context windows to 1M tokens without facing prohibitive decode-time memory growth. This design dramatically reduces memory bandwidth demands during extended multi-turn agent sessions.

NaiveAI claims the complete removal of full attention produces zero degradation on long-context retrieval while reducing prefill compute. Hardware testers counter that specialized sparse attention layouts like DSA require custom GPU kernels that lack widespread backend support outside of NaiveRT.

Verified across 1 sources: MindStudio (Sep 30)

SGLang PR #41855 Adds Opt-In CANN Sparse-Attention for Huawei Ascend on Qwen4-Exp

On Wednesday, September 30, 2026, SGLang pull request #41855 was submitted, introducing an opt-in CANN sparse-attention path for Huawei Ascend 910C hardware tailored to the unreleased Qwen4-Exp QSA architecture. The implementation builds a layout adapter that packs K, V, and Q matrices to invoke `torch_npu.npu_sparse_flash_attention` without modifying underlying model weights. The operator passes single-node correctness tests, though distributed tensor parallelism remains unvalidated.

Upstream pull requests for unreleased architectures like Qwen4-Exp provide early technical visibility into upcoming model structures. Furthermore, adding native sparse attention adapters for non-NVIDIA silicon expands hardware choices for serving large open-weight models on alternative accelerators.

The PR author confirms that packing tensor layouts on the fly enables native Ascend execution without weight conversion. SGLang reviewers note that full multi-node serving support cannot be merged until tensor-parallel communication collectives are fully implemented in CANN.

Verified across 1 sources: Orca Router (Sep 30)

Anthropic & Claude

Anthropic Ships Claude Code 2.1.286 Adding Permission Prompt Counters and Refusal Retries

Following yesterday's coverage of version 2.1.285, Anthropic released Claude Code version 2.1.286 on Wednesday, September 30, 2026. The CLI update introduces stacked permission prompt counters (e.g., '2 of 5'), automated fallback retries to previous model tiers upon encountering API refusals, and fixes for session resume failures caused by parallel tool calls. It also corrects pricing meter calculations on the Claude apps gateway, improves secret redaction for percent-encoded tokens, and restricts `--bare` mode to named MCP servers while suppressing background tasks.

Automated retries on API refusals and fixes for state corruption following parallel tool calls directly address workflow breaks in multi-step coding agent sessions. For developers relying on Claude Code for autonomous refactoring, stricter session recovery and secret redaction prevent context loss and accidental credential exposure in execution logs.

Anthropic positions the update as a stability release focused on enterprise safety and execution continuity. Developers using alternative proxy routers note that auto-retrying on refusals can unexpectedly burn API credits if an underlying system prompt continuously triggers safety filters.

Verified across 1 sources: AI TLDR (Sep 30)

Mechanistic Interpretability

etalii.dllm Tooling Suite Releases Deterministic SAE Training, Logit Lens, and Bit-Exact Tracing

On Thursday, October 1, 2026, maintainers updated the open-source `etalii.dllm` repository with Phase 11 features targeting mechanistic interpretability. The release adds `dllm sae train` for sparse autoencoders with deterministic gradient kernels that produce byte-identical files across identical seeds, `dllm lens` for total-order tie-breaking logit lens heatmaps in text, JSON, and HTML, and a `Transformer.trace()` Python API. The tracing module extracts residual streams, intermediate MLP activations, and head-level attention probabilities while matching original logits bit for bit.

Platform-dependent floating-point non-determinism frequently corrupts mechanistic interpretability experiments, making it difficult to reproduce feature probes across different GPU runs. By enforcing deterministic gradient kernels and bit-exact tracing APIs, `etalii.dllm` provides local researchers with a stable foundation to build custom activation patching and SAE probing toolkits.

Maintainers highlight that enforcing exact total-order tie-breaking on token IDs prevents false feature shifts during automated steering scans. Independent tool builders note, however, that disabling non-deterministic atomic operations in CUDA kernels trades off roughly 15-20% of SAE training throughput.

Verified across 3 sources: GitHub (Oct 1) · GitHub (Oct 1) · GitHub (Oct 1)

Agent Orchestration & Evals

A-Evolve Framework Automates Agent Optimization via File-System State Evolution

On Wednesday, September 30, 2026, A-EVO Lab open-sourced A-Evolve, a Python framework that applies evolutionary algorithms to optimize AI agent scaffolding. The tool treats system prompts, skills, and memory contexts as file-system states, wrapping around any agent via a minimal `solve()` method. Using automated mutation, validation, and git-tagged versioning, A-Evolve reached top ranks on MCP-Atlas and SWE-bench Verified while supporting Anthropic, OpenAI, and Bedrock backends.

Manual system-prompt tuning and static harness design fail to generalize across diverse coding and orchestration tasks. By automating prompt and skill mutations through git-tracked file-system states with automatic rollbacks, A-Evolve provides a reproducible methodology to discover higher-performing agent configurations without full model re-training.

The developers contend that file-system state evolution provides complete auditability for every prompt modification. Skeptics point out that unconstrained evolutionary loops risk overfitting to specific benchmark evaluation suites if evaluation tasks are not sufficiently diverse.

Verified across 1 sources: Bright Coding Blog (Sep 30)

Hugging Face Releases Agent Error Dataset with 50,228 Structured Error-Diagnosis Pairs

On Thursday, October 1, 2026, researchers published the Agent Error Dataset (AED), containing 50,228 structured error-diagnosis pairs compiled across 9,961 source tasks, 33 environments, and 19 harness families. Using a five-stage error-to-training pipeline, the authors trained models on failure trajectories rather than successful runs. On 3,062 matched replay pairs, first-proposal corrections increased verifier pass rates from 18.4% to 51.1%, while fine-tuning Qwen3-8B on diagnosis tasks raised step-agreement from 47.2% to 63.6%.

Post-training pipelines typically discard failed agent trajectories, missing crucial diagnostic data. Training models explicitly on structured error-diagnosis pairs equips agent frameworks to analyze intermediate execution failures, self-correct buggy tool calls, and improve verification pass rates during complex software engineering tasks.

The paper's authors emphasize that learning from failure trajectories produces far greater verifier gains than fine-tuning on positive trajectories alone. Evaluation researchers note that diagnostic accuracy depends heavily on whether the execution environment provides verbose, structured stack traces.

Verified across 1 sources: AI Weekly (Oct 1)

Local Inference Tooling

Magnitude Open-Sources Stateful Local Inference Engine for Autonomous Agents

On Thursday, October 1, 2026, engineers open-sourced Magnitude (Apache 2.0), a Rust-based inference runtime built specifically for agentic execution loops. Magnitude features on-device GPU kernel autotuning, dynamic session-based memory allocation, and hybrid paged attention. Benchmarked against `llama.cpp` running Qwen3.6-35B-A3B at 64k context, Magnitude demonstrated up to 92% faster decode speeds on Metal and 19% faster decode on CUDA while reducing per-agent memory consumption by 27–28%.

Traditional local inference engines treat agent tool calls as isolated, stateless requests, causing redundant prefill passes and locking up VRAM between execution turns. By fusing KV cache lifecycles directly with agent session states and autotuning GPU kernels locally, Magnitude allows practitioners to run multi-step agents on consumer hardware with significantly lower latency and VRAM overhead.

The authors argue that fusing session state into the runtime eliminates redundant prompt re-prefills across agent loops. System developers observe that on-device kernel compilation introduces a noticeable initial startup pause when first initializing a new model architecture.

Verified across 2 sources: Hacker News (Oct 1) · DEV Community (Oct 1)

MLX Outperforms llama.cpp on Short Prompts on Apple Silicon, but Lags at Long Context

A performance analysis published Wednesday, September 30, 2026, evaluated Apple's MLX framework against `llama.cpp` across Apple Silicon Macs. MLX demonstrated a 20–40% generation speed lead during short-prompt chat interactions. However, at context lengths between 30K and 146K tokens, `llama.cpp` using flash attention reversed the position, delivering roughly double the decode speed of MLX. The report also noted that Ollama has added an experimental MLX backend preview.

Local LLM practitioners cannot assume a single backend provides optimal performance across all workloads. While MLX leverages Apple Silicon's unified memory bandwidth efficiently for short turns, multi-step agent coding sessions carrying large context histories suffer severe regressions unless served through `llama.cpp`'s optimized flash attention kernels.

Benchmarkers highlight that `llama.cpp`'s C++ flash attention kernels handle long-context key-value lookups with significantly lower cache overhead than MLX's current Metal graphs. MLX advocates counter that upcoming framework releases will integrate fused long-context attention primitives that close the gap.

Verified across 1 sources: Vetted Consumer (Sep 30)

Maru Issue #95 Identifies vLLM Hybrid KV Cache Promotion Crash on Qwen3.5 Models

On Thursday, October 1, 2026, GitHub issue #95 was filed against the Maru repository, documenting an engine startup failure when serving Qwen3.5 hybrid models in vLLM. When `MaruKVConnector` is enabled, vLLM's Hybrid KV Cache Manager (HMA) is automatically disabled because the connector does not implement `SupportsHMA`, causing vLLM to throw a fatal promotion error on startup due to mismatched KV cache groups across Gated DeltaNet and full-attention layers.

As models combining Gated DeltaNet linear layers and full-attention layers become widely adopted, serving engines must manage heterogeneous KV cache allocations within the same process. This issue highlights ongoing integration friction between custom KV cache connectors and vLLM's internal hybrid memory management hooks.

The issue reporter provided a minimal reproduction script using `Qwen3.5-0.8B`, proving that custom connectors crash unless explicitly refactored to support HMA specifications. Engine maintainers are drafting an interface patch to expose multi-group KV specs cleanly to external connectors.

Verified across 1 sources: GitHub (Oct 1)

Quantization & KV-Cache

REAL-Q Framework Introduces Dynamic Block-Wise Gradient Descent for LLM Quantization

A paper published Wednesday, September 30, 2026, presented REAL-Q (End-to-End Quantization via Dynamic Gradient Descent), addressing structural flaws in GPTQ such as static Hessian approximations. REAL-Q uses an aggregated Fisher matrix to approximate global KL divergence and runs a dynamic block-wise gradient descent step using Adam every 128-column block. Tested on LLaMA-3.1 and Qwen3 checkpoints at W4A16 precision, REAL-Q achieved up to a 49% reduction in end-to-end KL divergence compared to GuidedQuant.

Traditional layer-local post-training quantization solvers like GPTQ ignore how quantization noise accumulates across sequential layers. By optimizing blocks against a global KL divergence surrogate during quantization, REAL-Q preserves generation quality in 4-bit models without requiring expensive full-network fine-tuning.

The authors show that dynamic Adam updates per column block prevent early-layer quantization errors from ruining late-layer representations. Inference engineers point out that running iterative Adam steps during quantization increases PTQ processing times significantly compared to standard one-shot GPTQ.

Verified across 1 sources: Dev.to (Sep 30)

ML Systems & Hardware

GMKtec Ships Evo-X5 Pro Featuring AMD Gorgon Halo and 192GB Unified LPDDR5X Memory

On Monday, September 28, 2026 (analyzed September 30), GMKtec released the Evo-X5 Pro mini PC, starting at $6,799. It is the first commercial system shipping with AMD's Ryzen AI Max+ Pro 495 'Gorgon Halo' processor and 192GB of LPDDR5X-8533 unified memory across a 256-bit bus, delivering 273 GB/s of bandwidth. The memory capacity allows local practitioners to load 300B-parameter models at 4-bit quantization without NVMe streaming or multi-node clustering.

For local LLM practitioners, unified memory capacity dictates whether massive open-weight models can run locally at all. A 192GB unified memory pool on a desktop system provides a middle ground between consumer GPUs constrained to 24GB/32GB VRAM and expensive multi-GPU server setups, enabling full-model residency for high-parameter MoE architectures.

Hardware analysts note that while 192GB of unified RAM removes model offloading bottlenecks, 273 GB/s of memory bandwidth limits decoding speed to roughly 3–5 tokens/sec on 300B models. Enterprise buyers also flag jurisdictional data compliance considerations for hardware produced by Shenzhen-headquartered manufacturers.

Verified across 1 sources: Tech Times (Sep 30)

GitHub Feature Request Proposes MLX Backend for Decider Models on Apple Silicon

A feature request (Issue #19) filed on the Decider repository on Thursday, October 1, 2026, requested a native MLX inference backend for dense decision models like `decider-4b` running on Apple Silicon. The author noted that while existing paths rely on PyTorch MPS with optional Metal kernels for gated-delta operations, a pure MLX runtime would eliminate PyTorch dependency overhead, enable resident BF16/4-bit decision services, and improve compatibility with Mac agent harnesses.

Running small, resident decision-making models alongside local agent harnesses on Apple Silicon requires minimizing runtime memory footprints. Porting PyTorch MPS decision models to native MLX backends reduces memory residency, allowing background decision routers to remain warm without hogging unified RAM.

The issue author points out that PyTorch MPS allocations create unnecessary VRAM overhead for simple classification tasks. Maintainers agree in principle but note that custom Gated DeltaNet operations must be re-written in MLX C++ extensions to maintain parity.

Verified across 1 sources: GitHub (Oct 1)

Open-Weights Policy

GitHub H1 2026 Data Highlights State AI Legislation and Open Source Exemptions

On Wednesday, September 30, 2026, GitHub published its H1 2026 Transparency Center report. Alongside tracking government takedown requests and DMCA Section 1201 rulemaking, the report detailed legislative engagement across U.S. states. Notably, it highlighted efforts to align California's AI Transparency Act (SB 1000) with open-source licenses by replacing broad compliance mandates with a narrower notice-and-response framework for developer tool repositories.

State-level AI regulations can inadvertently create legal liability for developer platforms hosting open-weight models and developer tooling. Securing explicit open-source license exemptions in legislation like California's SB 1000 ensures that independent developers can continue downloading, hosting, and modifying open models without facing institutional compliance burdens.

GitHub advocates emphasize that notice-and-response frameworks protect open-source maintainers from liability for downstream model uses. Policy analysts counter that state-by-state legislative variations still create a fragmented compliance landscape for open-source AI distribution.

Verified across 1 sources: TechGig (Sep 30)


The Big Picture

Recurrent State Quantization Focuses on Memory Lifespan Frameworks like STEPQuant and LeapQuant demonstrate that linear attention recurrent states cannot be treated like standard KV caches under post-training quantization. By introducing time-space precision allocations and compensator tokens, engineers are mitigating state error accumulation during autoregressive generation.

Local Serving Engines Autotune for Long-Context Memory Ratios Engines including SGLang v0.5.19, Magnitude, and vLLM are shipping dedicated memory allocation flags and custom paged memory pools specifically tailored for hybrid models like Qwen3.8-27B and Decider-2B that mix linear-recurrent layers with standard full attention.

Deterministic Tooling and Byte-Identical Probes Standardize Interoperability Toolkits like etalii.dllm and agent platforms like A-Evolve are prioritizing bit-exact tracing, seeded SAE training, and file-system state versioning to make interpretability experiments and agent self-optimization entirely reproducible.

Massive Unified Memory Desktops Challenge Offloading Architectures Configurations with 192GB of unified LPDDR5X memory, such as AMD Gorgon Halo systems, allow practitioners to host 300B-parameter sparse models locally without relying on high-latency NVMe streaming or multi-node tensor parallelism.

Subagent Orchestration Standardizes Outer-Loop Verification Gates Agent harnesses are shedding monolithic prompt engineering in favor of structured parent-child task isolation, strict fallback file schemas like AGENTS.md, and automated checker loops that enforce default-FAIL evidence contracts.

What to Expect

2026-10-15 — StepFun expected open-weight release of Step 5 600B MoE foundation model.
2026-11-01 — Scheduled maintenance update for SGLang Ascend CANN tensor parallelism extensions.

Every story, researched.

Every story verified across multiple sources before publication.

🔍

Scanned

Across multiple search engines and news databases

426
📖

Read in full

Every article opened, read, and evaluated

107
⭐

Published today

Ranked by importance and verified across sources

19

— The Bandwidth-Bound

🎙 Listen as a podcast

Subscribe in your favorite podcast app to get each new briefing delivered automatically as audio.

Apple Podcasts
Library tab → ••• menu → Follow a Show by URL → paste
Overcast
+ button → Add URL → paste
Pocket Casts
Search bar → paste URL
Castro, AntennaPod, Podcast Addict, Castbox, Podverse, Fountain
Look for Add by URL or paste into search

Spotify isn’t supported yet — it only lists shows from its own directory. Let us know if you need it there.