<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Signal Over Noise]]></title><description><![CDATA[Most AI content is written by people selling something. This isn't. Signal Over Noise is a practitioner-first publication on AI agents and automation, written by a CTO who's made the expensive mistakes so you don't have to.]]></description><link>https://signalovernoise.tech</link><image><url>https://substackcdn.com/image/fetch/$s_!Owu3!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb7415cf-9ce5-438f-98c5-7acddf07ed8f_1280x1280.png</url><title>Signal Over Noise</title><link>https://signalovernoise.tech</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 02:29:03 GMT</lastBuildDate><atom:link href="https://signalovernoise.tech/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Justin Wilson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[signalovernoisetech@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[signalovernoisetech@substack.com]]></itunes:email><itunes:name><![CDATA[Justin Wilson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Justin Wilson]]></itunes:author><googleplay:owner><![CDATA[signalovernoisetech@substack.com]]></googleplay:owner><googleplay:email><![CDATA[signalovernoisetech@substack.com]]></googleplay:email><googleplay:author><![CDATA[Justin Wilson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Model Does Not Have to Fit. The Working Set Does.]]></title><description><![CDATA[Loading is a screenshot.]]></description><link>https://signalovernoise.tech/p/the-model-does-not-have-to-fit-the</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-model-does-not-have-to-fit-the</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Tue, 01 Sep 2026 09:48:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UWWg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Loading is a screenshot. A workflow needs latency you can live with and quality that survives the packing.</em></p><div><hr></div><p>A screenshot of a loaded model is not a workflow. nvidia-smi at 22 GB tells you the weights landed. It does not tell you whether the second agent turn still has a tolerable time to first token, whether a 100k prompt still retrieves, or whether decode is fast enough that you stay in the thought.</p><p>August spent thirty-one days on whether the agent is still there after you close the terminal. The model underneath that agent has a quieter failure. The process is up. The conversation is waiting. Local inference culture treats that gap as a VRAM number. Discord asks &#8220;does it fit?&#8221; and a green memory bar answers. Local AI spent the last year celebrating packing victories that nobody would use twice. Fit is the wrong first question. The working set is the right one.</p><p>Qwen3.8-27B at Unsloth UD-Q4_K_M is 15.33 GB on disk. On a 24 GB RTX 3090 that looks like headroom. I measured the rest of the bill on August 22. The same file at 64k context with FP16 KV and MTP already sat around 21.8 GB. Pushing the context to 131,072 with FP16 KV is the OOM path. Q8 KV and a single slot made 128k fit at 22.5 GB. Same weights. Different working set. The context length printed on the model card did not make 128k free.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UWWg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UWWg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UWWg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1463892,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/213680253?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UWWg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UWWg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffdd2a397-094d-42f0-8d6d-ca381eae0411_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The working set is whatever has to be hot for the next operation. Dense weights, or the experts the router actually picked. The KV cache for the live context. Activations. In an image graph, the text encoder, the VAE, and the latents that are in flight. Those pieces do not all need the same pool at the same time, and they do not all need to be resident for the whole run. Treating them as one blob is how a &#8220;fits&#8221; screenshot becomes a forty-second prefill and a decode that feels like a demo.</p><p>An agent makes this worse, because the expensive phase is often the one people stop measuring. Prefill is prompt processing. Decode is generation. A chat demo is mostly decode. A tool-using agent edits the context after every call. If the prefix cannot be reused, you pay prefill again. On that 3090, a 99,226-token prompt took 131 seconds to prefill at 761 tok/s. An agent that does that on every turn is not local intelligence. It is a wait with a model attached.</p><p>People screenshot nvidia-smi because it is the only number the stack makes easy. llama-server reports slots and context, not whether the last tool turn reused the prefix. Ollama reports that the model is running, not the KV type that made 128k possible. Local tooling still reports capacity. Capacity is the least interesting tier once the process starts.</p><p>The obvious answers look like they dissolve the constraint. They mostly move it.</p><p>llama.cpp will offload leftover layers into system RAM. That is a real feature, and it is also a bandwidth tax. Every token that hits a CPU-resident layer pays host memory and the PCIe pipe instead of staying on the card. I already have the ugly version of that sentence on this hardware. gpt-oss-120b is about 46 GB at every GGUF quant I tried, and offload is the 5 tok/s path. The process starts. Nobody would pair it with a runtime that expects a reply.</p><p>Mixture of Experts makes the same mistake in marketing copy. &#8220;3B active&#8221; sounds like a 35B model occupies 3B of memory. Active parameters are compute per token. The expert pool still has to live somewhere. Decode may touch a small routed set. Long-prompt prefill can touch nearly every expert. Colibr&#236; is the honest version of that idea, not the slogan. GLM-5.3-Flash is 321B total and can sit around 12 GB resident. The project&#8217;s own reference SSD still prices decode at roughly 44 seconds per token cold, because one token moves about 4.8 GB of experts. That is a systems result worth studying. It is not a daily driver. Loading is not success.</p><p>Two cards invite the same arithmetic error. A 16 GB card next to a 24 GB card is not a 40 GB GPU. Capacity can be partitioned. Transfer still happens. Apple unified memory removes the discrete VRAM wall and still has a capacity budget and a bandwidth budget. The number on the product page is not the working set, and adding two product pages does not merge the pools.</p><p>The useful frame is four tiers, not one number.</p><p>VRAM, or the GPU side of unified memory, is the hot pool. Compute lives here. If the working set for the current phase fits here, decode can be interactive. System RAM, or the host side of unified memory, is the warm pool. Weights can live here. Experts can live here. The tax is the pipe into the compute unit. PCIe is not storage. It is a clocked pipe. Every layer, expert, or KV page that crosses it costs latency you cannot recover in a better quant. NVMe is the cold pool. Treating the SSD as an expert store is a legitimate architecture. It is also why &#8220;runs in 25 GB of RAM&#8221; and &#8220;usable chat&#8221; are different sentences.</p><p>On the 3090 run, the 15.33 GB of Q4 weights and the Q8 KV lived in VRAM. RAM held the OS and llama.cpp&#8217;s host-side bookkeeping. PCIe stayed quiet because nothing was spilling. NVMe was involved once, at load. That is why 99 tok/s was possible. The gpt-oss-120b offload run inverted the map: most of the weights sat in RAM, every token crossed PCIe, and 5 tok/s was the honest number. Same machine. Different placement. Bandwidth is the part the screenshot never shows. Once the working set spills out of the hot pool, you are no longer running at GPU memory speed. You are running at whatever pipe you spilled onto. A clever placement policy can hide that on decode if the hot slice is small. Prefill, long context, and a cache miss will not let you hide it. The engine that looks magical on a short prompt is often the engine that falls over on the second agent turn.</p><p>Concurrency is part of the working set too. The 64k llama-server I ran advertised four slots. VRAM will not hold four concurrent 64k fills. One user at 128k is a different machine than four users at 8k. Desktop VRAM is also not a dedicated inference allocation. The OS, a compositor, and whatever else is using the card already took a cut. &#8220;Fits in 24 GB&#8221; always meant &#8220;fits in whatever is left.&#8221; The same hierarchy returns in image and video graphs later this month. A denoiser file that fits is still not a workflow if the VAE and the latent at your resolution do not.</p><p>The test for the next thirty days is three questions, and a recommendation has to answer all three. Fit: does the working set for this phase land in a pool that can feed the compute? Speed: is time to first token and decode tolerable for the job, including the second turn? Quality: did the packing, the KV type, and the offload preserve retrieval and answers you would actually ship?</p><p>I ran that test on the 3090. Qwen3.8-27B UD-Q4_K_M, 128k context, Q8 KV, one slot, MTP draft 5: 22,541 MiB used, 99 tok/s decode, 761 tok/s prefill on that 99k prompt, and a needle buried at the midpoint that came back as ORANGE-PLUM-917 exactly. Drop MTP and decode halves to 42 tok/s. Same file. Same card. The working set changed. Q4 KV saved about 2 GB at the same draft-2 speed. I have not treated that cache type as proven for the 100k retrieval. Memory you saved in the cache is memory you may have spent in accuracy. Unsloth UD-Q6_K is 20.47 GB on disk and will not fit the 64k-plus-MTP layout I actually ran. The quality ceiling on 24 GB is not the highest quant filename. It is the highest quant that still leaves room for the cache and the draft. Fit without the quality gate is packing.</p><p>llama.cpp is the boring baseline every exotic claim this month has to beat. Thursday is that tour: GGUF, partial offload, quantized KV, MTP, the OpenAI-compatible server, and why the first run still belongs there even when a specialized engine may eventually win. Tomorrow is the map. Where weights, KV, activations, experts, and checkpoints can live, what moves during prefill versus decode, and how a discrete GPU box and a unified-memory box draw the same hierarchy with different labels. Friday collects the claims that only proved fit and the ones that proved a usable workflow.</p><p>The rest of September is engines that move the working set on purpose. Colibr&#236; streams routed experts from disk. FreeToken treats the whole PC as one elastic system. KTransformers puts dense work on GPU and experts where they belong. MoE-Infinity treats expert caching as a serving problem. Later weeks pick a runtime for the box you own, spend the leftover bytes on KV and speculative decoding, then watch the same hierarchy show up in image and video graphs. None of that cancels physics. Each engine chooses which tier pays.</p><p>August asked whether the agent survives the terminal. The model underneath that agent has the same problem one layer down. A runtime that cannot get a token back in time looks idle. Pick the strongest model whose working set meets the latency and quality floor of the job. The largest model that prints one token is a screenshot.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[What August 2026 Told Us About the Agent Execution Layer]]></title><description><![CDATA[Five arcs, one split: frameworks still define the agent.]]></description><link>https://signalovernoise.tech/p/what-august-2026-told-us-about-the</link><guid isPermaLink="false">https://signalovernoise.tech/p/what-august-2026-told-us-about-the</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Mon, 31 Aug 2026 10:19:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uMBW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Five arcs, one split: frameworks still define the agent. Runtimes decide whether it is still there in the morning.</em></p><div><hr></div><p>August asked one question for thirty-one days: does the agent keep running after you close the terminal? I already knew the polite answer. A demo survives a meeting. A runtime survives a crash, a credential rotation, and a weekend with nobody watching the process. The month was not about discovering that distinction. It was about watching the field keep selling definition tools as if they were operators.</p><p>The opening arc made the boundary expensive. A framework gives you nodes, roles, and a launch method. LangGraph will checkpoint a graph to SQLite or Postgres. CrewAI Flows will persist a Pydantic state object across kickoff(). AutoGen will still show a crowd for a tree whose last feature tag is from September 30, 2025, with a last push on April 15, 2026. None of that owns the loop after the worker dies. Hermes Agent and OpenClaw do. Hermes, tagged v2026.8.27 as v0.20.6 on August 27, treats skills as accumulated procedure with trust scores, durable sessions, and a delivery-obligation ledger so a finished answer cannot vanish between generation and send. OpenClaw treats every message, cron tick, webhook, and heartbeat as an event on one gateway, which is why channel adapters are the product and why the repo sat at 388,041 stars by yesterday&#8217;s scorecard. Restate sits next to both as the primitive most homemade loops skip: resume from the last completed step, not from the prompt. The CASE that closed the week was a loop, not a library: event queue, context assembly, tool discovery, model call, action dispatch, checkpoint. Skip the last box and you have a launch script.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uMBW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uMBW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uMBW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1439918,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/213523934?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uMBW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!uMBW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc01d4f72-e182-48cc-8980-6c1af8e96344_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The 98.4 percent finding from MBZUAI kept returning because it is the same sentence in a lab coat. Four independent agents reverse-engineered Claude Code and found 98.4 percent harness: permissions, context, sandboxing, tool routing, recovery. 1.6 percent decision logic. The industry still spends its attention in the inverse ratio. That is not a model problem. That is a budget problem wearing a research paper.</p><p>The desktop arc moved the same argument onto a machine you own. A cloud session resets. A desktop agent accumulates. Goose, Block&#8217;s open-source companion, and BeeAI Desktop, the Linux Foundation&#8217;s interop bet, are not competing chat apps. They are bets that the agent&#8217;s state belongs next to your files, your credentials, and your cron. The privacy case is real, and it is the weaker one. The compounding case is the one that changes the architecture: skills, memory, and tool paths that survive the night. Inference can live anywhere. State that wipes every morning is a product that cannot get better. Claude Code remains the reference harness, not because Anthropic invented agents, but because most of the product is the part vendors keep calling plumbing. The personal stack CASE at mid-month was the same loop as week one, with the process pinned to your hardware. Call the model wherever you want. Keep the context at home.</p><p>The multi-agent week was the week teams add a second agent and call it architecture. Supervisor works: one dispatcher, specialists with schemas, no debate. Blackboard works: agents write a structured workspace, not a transcript. Hierarchical works when the work is actually a tree, a FOIA fan-out across departments, not a single clinical judgment wearing extra titles. Swarm photographs well and produces invoices. CrewAI still wins the role-based demo, and Flows is the control plane that makes those roles replayable. Microsoft Agent Framework productized handoff, sequential, concurrent, GroupChat, and Magentic; GroupChat stays off a case file. DSPy 3.3 treated coordination as a compiler problem, which is the honest version of &#8220;the prompt is not the architecture.&#8221; The TAKE that earned the week was the refusal. Multi-agent is the wrong answer when the work does not split. Token inflation with extra names is not a pattern. It is a meeting.</p><p>The middleware week left the loop and asked how the agent touches a system you did not write. MCP is the USB-C of that question, and the July 28 spec made it small on purpose: no session, no handshake, one HTTP call, a schema. Composio is the registry bet. Strands Agents SDK is AWS&#8217;s quiet layer for the same job, framework-agnostic even when Bedrock is in the room. OpenConnector is the enterprise objection: not whether the agent can find the tool, but whether it should, and whether you can prove it. Authentication is still the hole. Catalogs discover. Hooks decide whether a call fires. Neither one can tell you who the agent is when the request leaves the process. Friday&#8217;s roundup made that concrete. The MCP roadmap named workload identity and token exchange as a spec gap, not a vendor feature. Patrick Walsh at IronCore Labs showed OpenClaw&#8217;s spotlighting tags dying at memory promotion, so an email that failed as an injection succeeded as a MEMORY.md fact. NVD scored a 10.0 against an MCP HTTP server that bound :: with auth off. USB-C with the port open and the lock optional is a LAN shell.</p><p>Yesterday&#8217;s scorecard was the decision rule, not a ranking. If the agent is the long-lived worker, pick a runtime. Hermes when that worker should get better. OpenClaw when it has to exist on the channels your users already live in, and you can live with a stable tag that lags a hot beta. If the workflow is the product, pick a framework and operate it yourself. LangGraph when the topology is a graph you must checkpoint and interrupt. CrewAI Flows when the team thinks in roles and the path has to replay for an auditor. Microsoft Agent Framework when procurement and Azure identity are the door. AutoGen is not on the list for new work. Saturday&#8217;s mixed path still holds: runtime at the edge for judgment and continuity, a service you own for authority. Flexibility is a clean ownership line, not a fear of dependencies.</p><p>Put the five arcs side by side and the month is a bifurcation, not a feature matrix. Frameworks define what the agent is. Runtimes determine whether it is still there after the process dies, the credential rotates, and the model call fails on step thirty-seven. The tools that will still matter in a year treat persistence, recovery, and skill accumulation as the product. The tools that will not treat those as an exercise for the reader. Star counts will keep lying. OpenClaw out-stars Hermes. AutoGen out-stars Agent Framework by more than four to one. The living tree is the one still tagging releases and still arguing about delivery after a gateway crash.</p><p>The surprise of the month was not a new runtime. It was how fast the scarce layer moved from discovery to identity. MCP can find the tool. Composio can catalog it. The call still leaves the process as an anonymous bearer token, and the memory file still swallows untrusted text once the spotlighting tags fall off. Persistence without provenance is how a runtime becomes an attack surface.</p><p>Hermes without skill review is drift with a flattering name. OpenClaw with business rules in MEMORY.md is a system of record you cannot replay. A graph you never made idempotent is a double email at 2:13 a.m. The month did not pick a winner. It named the failure each choice buys.</p><p>September goes one layer down. The runtime still depends on the model sitting under it, and &#8220;the model&#8221; is no longer a single resident blob in VRAM. Loading is not success. Usable latency and preserved quality are. VRAM, system RAM, unified memory, PCIe, and NVMe are different tiers, not one interchangeable number. Dense weights, routed experts, KV cache, and activations do not all need the same pool at the same time. A screenshot of a loaded model is not a workflow. llama.cpp remains the boring baseline every exotic claim has to beat. The model does not have to fit. The working set does.</p><p>August&#8217;s question was whether the agent survives the terminal. September&#8217;s question is whether the box you already own can make the model underneath that agent useful. The gap between those two questions is where the next thirty days of this field get decided.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Agent Runtime Scorecard]]></title><description><![CDATA[Hermes, OpenClaw, LangGraph, and the Field]]></description><link>https://signalovernoise.tech/p/the-agent-runtime-scorecard</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-agent-runtime-scorecard</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sun, 30 Aug 2026 08:54:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!r8B_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>If you can only pick one, pick the failure you are willing to own.</em></p><div><hr></div><p>Star counts will send you to the wrong production. OpenClaw sits at 388,041 GitHub stars this morning. Hermes Agent sits at 238,242. AutoGen still shows 60,698, and its last feature tag is python-v0.7.5 from September 30, 2025. The successor, Microsoft Agent Framework, has 13,214 stars and shipped Python 1.16.0 two days ago. Rank the field by crowd size and you will adopt a museum while skipping the repo that is actually moving.</p><p>Yesterday I mapped three paths: build the runtime around a framework, adopt OpenClaw for reach, or bet on Hermes for compounding. Those paths still hold. The scorecard is how you test them against the failures that show up after the demo. Persistence, recovery, memory, tool discovery, multi-agent support, community, self-improvement, token spend, and the cost of owning the thing. Feature matrices flatten those into checkmarks. Production does not.</p><p>Five names keep landing in the same meeting, and they are not the same kind of product. Hermes Agent v2026.8.27, tagged August 27 as v0.20.6, is a Python runtime you operate. OpenClaw&#8217;s npm latest and GitHub &#8220;latest&#8221; tag are v2026.7.1-2 from August 4; the tree is also shipping v2026.9.1-beta.1 as of August 28. LangGraph is 1.2.11 on PyPI, released August 11, at 40,692 stars. CrewAI is 1.15.18, released August 27, at 57,816. Microsoft Agent Framework is Python 1.16.0 and .NET 1.19.0. They can all run useful agent work this week. They do not own the same loop.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!r8B_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!r8B_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!r8B_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1389103,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/213377403?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!r8B_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!r8B_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb962c211-f4d0-4d7b-b05c-248b36037f03_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>What they agree on is narrower than the marketing. Every one of them will persist something. Hermes writes sessions, skills, memory, and checkpoints into durable state with rollback. OpenClaw keeps memory files, skills, and instructions in a workspace, and the last month of releases added crash-recoverable SQLite snapshots and quarantine stores. LangGraph will checkpoint graph state to SQLite or Postgres. CrewAI Flows will persist a Pydantic state object across kickoff(). Agent Framework 1.16.0 even had to stop workflow checkpoints from being mutated outside storage, which tells you the persistence is real enough to grow bugs. AutoGen is the control: it no longer ships this layer at all.</p><p>Agreement ends at the process boundary. A checkpointer saves the graph. It does not notice the dead worker, replay a side effect safely, or deliver the last message after the gateway dies between generation and send. That is the runtime question August has been asking since the first Saturday. LangGraph, CrewAI, and Agent Framework give you graph-level resume inside a service you already run. Hermes and OpenClaw give you an operator for an agent that is supposed to exist when nobody is watching. Mix those two rows on a spreadsheet and you will buy a library while expecting a pager.</p><p>Recovery is where the split gets expensive. Hermes already shipped a delivery-obligation ledger so a finished response cannot vanish if the gateway crashes between generation and delivery. The August 27 window added durable incident acknowledgements for cron and clearer code-skew failures. OpenClaw&#8217;s current beta is still landing gateway restart recovery: preserve admitted turns across repeated restarts so a restart-safe run continues through each checkpoint and delivers the final response. That is the right problem. It is also still a beta line, while the stable npm package remains 2026.7.1-2. LangGraph resume is excellent if you own the worker fleet that wakes the graph. CrewAI will resume a flow you persisted. Agent Framework will resume a workflow you hosted. None of the frameworks will page you when the process that hosts them is gone.</p><p>Memory is the dimension most scorecards fake. Hermes scores memories, accumulates procedural skills, and carries those procedures into a cron run, a terminal session, or a Slack message. That compounding loop has not changed in the month since I wrote about it. The surface around it has: lean-tail compression is now the default, tool search is multi-query with stemming, and the remote MCP catalog added fifty-plus live-verified vendor-hosted servers in the v0.20.6 window. OpenClaw has memory files and skills. It does not put an autonomous closed learning loop at the center of the product. LangGraph, CrewAI, and Agent Framework will use whatever memory you wire. That is not a defect. It is a statement that memory is your problem, which is the correct statement for a library.</p><p>Tool discovery follows the same split. Hermes and OpenClaw treat tools as something the runtime can find while it is running: MCP servers, plugins, ClawHub, a catalog, a search path. Frameworks treat tools as functions you bind in code. After a week on MCP, Composio, Strands, and OpenConnector, that difference should feel familiar. If the agent&#8217;s job is to live on a changing set of APIs, runtime discovery is the product. If the agent&#8217;s job is a known graph of approved actions, binding the tools in code is the control plane. Skip Hermes for the catalog size when your auditor needs a frozen allowlist. Skip LangGraph for the frozen allowlist when your users keep adding channels.</p><p>Multi-agent support is the trap for personal runtimes. CrewAI still wins the role-based demo, and Flows is the control plane under it. Agent Framework productized handoff, sequential, concurrent, GroupChat, and Magentic, and I already said GroupChat stays off a case-file workload. LangGraph will express multiple actors as a graph you can checkpoint. Hermes has profiles and subagents inside one runtime. OpenClaw is a personal agent surface that can be extended, not a coordination fabric. Start with a framework inside a service when the work is a crew. Start with a runtime when the work is one agent with isolated identities. Borrow a crew abstraction from a personal assistant and you get token inflation with extra names.</p><p>Community size is the number everyone quotes and the number that lies with the most confidence. OpenClaw has the crowd, and that crowd matters for channel adapters that rot. Hermes has a quarter-million stars and a release cadence that rolled about 525 pull requests into v0.20.6 in six days. CrewAI ships near-daily patches. LangGraph is slower on the core package, 1.2.11 on August 11, and busier on the SDK and checkpointer satellites. AutoGen still out-stars Agent Framework by more than four to one, with a last push on April 15. Stars buy you shared maintenance on undifferentiated glue. They do not buy you a living tree. Use them to estimate who will see a Slack scope change first. Do not use them to estimate who will still be tagging releases in six months.</p><p>Self-improvement is a one-horse dimension. Hermes has the loop. OpenClaw does not, by architectural choice. The frameworks do not, because they are not operators. If your value comes from procedures that accumulate across incidents, this dimension is the decision. If your value comes from a graph that must replay the same way for an auditor, this dimension is a distraction. I would rather have a slow pipeline I trust than a skill library I cannot review. Hermes without skill review is drift with a flattering name. That cost belongs on the scorecard next to the capability.</p><p>Token efficiency is where multi-agent theater goes to die. A CrewAI crew or an Agent Framework GroupChat pays a full prompt for every speaker, every round. LangGraph is as cheap as the graph you wrote. OpenClaw sessions grow like any long-lived chat unless you budget the prompt. Hermes now defaults to lean-tail compression, and a skill that already knows your verification commands stops spending tokens rediscovering the repo. Hermes does not always win this column. The runtime that reuses procedure beats the framework that restages a meeting. A graph of three deterministic nodes and a human interrupt will beat both personal runtimes on spend, because there is no extra operator in the way.</p><p>Total cost of ownership is the dimension that decides the meeting after everyone is tired of features. A framework looks cheap: pip install langgraph, or crewai, or agent-framework, on top of a service, an identity provider, and an on-call rotation you already have. The hidden cost is every concern outside the graph. A runtime looks expensive: you are standing up Hermes or OpenClaw as a platform, with profiles, gateway config, channel auth, and a learning curve. The hidden savings is the year of glue you do not write for messaging, scheduling, session storage, and crash recovery. Count engineering time over twelve months, not install time over twelve minutes. Count who is awake when the process dies.</p><p>The deciding factor is still who owns the loop. If the agent is the long-lived worker, pick a runtime. Hermes is the bet when that worker should get better. OpenClaw is the bet when it must exist on the channels your users already live in, and you can accept a stable tag that lags a hot beta line. If the workflow is the product, pick a framework and operate it yourself. LangGraph earns that seat when the topology is a graph you need to checkpoint and interrupt. CrewAI Flows earns it when the team thinks in roles and the path has to replay for an auditor. Walk through Microsoft Agent Framework when procurement, OpenTelemetry, and Azure identity are the door you cannot skip. AutoGen is not on that list for new work.</p><p>If you can only pick one, pick the layer that matches the 2:13 a.m. failure. A dead gateway with an undelivered answer is a Hermes or OpenClaw problem. A graph that woke twice and emailed the member twice is a LangGraph or CrewAI problem you failed to make idempotent. Import GroupChat and you will spend four hundred thousand tokens watching specialists agree with a hallucination; that one is on you, not on Microsoft. A ranking does not fix a wrong owner. It only tells you which owner you actually hired.</p><p>Tomorrow I will close August with what the month taught us about the execution layer, and hand September a problem one level down: the runtime still depends on the model sitting under it.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Build Your Own Runtime, Adopt OpenClaw, or Bet on Hermes]]></title><description><![CDATA[The Decision Framework]]></description><link>https://signalovernoise.tech/p/build-your-own-runtime-adopt-openclaw</link><guid isPermaLink="false">https://signalovernoise.tech/p/build-your-own-runtime-adopt-openclaw</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sat, 29 Aug 2026 08:52:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_00D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The choice is not build versus buy. It is which failure boundary you are willing to own.</em></p><div><hr></div><p>You can write an agent loop in an afternoon. You can spend the next year discovering that the loop was the cheapest part.</p><p>The expensive work begins when the model times out after a tool changed state, the process restarts halfway through a workflow, a channel delivers the same event twice, or a credential reaches a plugin that never needed it. None of those failures care whether the first prototype used LangGraph, CrewAI, OpenClaw, or Hermes Agent. They care who owns recovery, state, identity, and the next attempt. Your runtime choice assigns that ownership, whether the architecture document admits it or not.</p><p>August has spent four weeks separating the framework from the runtime. LangGraph defines a graph and can checkpoint its state. CrewAI Flows gives a Python workflow explicit starts, listeners, routers, and persisted execution. OpenClaw owns an integrated agent surface across model discovery, tools, sessions, and channel delivery. Hermes owns the conversation loop, skills, memory, scheduling, profiles, tools, and messaging gateway. All four can run useful agent work. They are not interchangeable, and treating them as a feature checklist produces the wrong decision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_00D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_00D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_00D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_00D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_00D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_00D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1438941,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/213254590?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_00D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_00D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_00D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_00D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F840facaa-af17-46d9-b5bb-3cc7ffae3517_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The first path is to build the runtime around a framework. Teams usually describe this as maximum control, which is true in the least useful sense. You control every queue, schema migration, retry policy, approval gate, credential boundary, deployment, and incident. LangGraph can save graph state and pause on an interrupt. CrewAI can persist a flow and resume it. Your platform still has to notice the dead worker, wake the run, decide whether the last side effect can be repeated, enforce tenant boundaries, rotate credentials, and explain the sequence to an auditor. A checkpointer is part of a runtime. It is not the operator.</p><p>That ownership is worth paying for when the workflow is the product. A prior-authorization engine, claims adjudication path, or government case-routing system should not hide its decision topology inside a general-purpose personal agent. The branch from eligibility to clinical review belongs in code, with state captured in a schema and human approval at a named transition. Audit events should land in a store your compliance team already knows how to query. LangGraph or CrewAI Flows can make that control plane easier to express, but the surrounding runtime remains yours because the domain boundary is yours.</p><p>The hidden cost is not the graph. It is every concern outside the graph that becomes visible after launch. Idempotency for tool calls needs a definition before production. When a queue wakes five hundred runs after an outage, the concurrency cap is yours. Replay must avoid sending the same email twice, checkpoint migrations must survive workflow changes, and support will outlive the engineer who wrote the first node. Build when those decisions contain business value. Building them to avoid adopting a runtime is expensive pride.</p><p>OpenClaw is the second path, and its strongest argument is not that it saves code. Its strongest argument is that it has already chosen an operating model.</p><p>OpenClaw treats the gateway as the center of the agent. Model discovery, tool wiring, prompt assembly, session management, and channel delivery live on one integrated surface. Its workspace keeps memory files, skills, and agent instructions together. The project is TypeScript-first and built for an agent that needs to exist across the channels where people already work. If the job begins with a WhatsApp message, continues in Slack, wakes on a schedule, and returns through the same gateway, that architectural opinion removes months of glue.</p><p>Community matters here because channel integrations decay. Slack changes scopes. WhatsApp changes onboarding. Mobile companions break on operating-system updates. A broad contributor base gives an operator a better chance that someone else sees the break first, writes the adapter, and documents the ugly edge. The benefit is not a star count. It is shared maintenance on the least differentiated part of an always-on assistant.</p><p>That bargain has a boundary. OpenClaw has memory and skills, but it does not put an autonomous closed learning loop at the center of the product the way Hermes does. Its architectural strength is connectivity: one agent surface, many channels, one event-shaped operating model. Teams that need a TypeScript-native assistant with broad reach should not dismiss that as shallow. Connectivity is the product when the agent fails if the user has to open another application. The mistake is expecting a channel runtime to become your domain workflow engine because plugins are easy to add.</p><p>Plugin convenience can turn into platform coupling one small decision at a time. A custom approval rule lands in a hook. A business transaction lands in a skill. A regulated decision lands in a memory file because the agent needs it tomorrow. Six months later, the runtime workspace contains business logic nobody can replay without the model. OpenClaw is the right choice when the gateway and channel surface are strategic. Keep the authoritative domain decisions behind explicit APIs, even when the plugin could do more.</p><p>Hermes is the third path. Its bet is not connectivity alone, though the gateway spans more than twenty messaging platforms. Its bet is compounding.</p><p>Hermes has a closed learning loop built around persistent memory and procedural skills. It can create a skill from a solved problem, improve that skill during later use, recall prior sessions through its session store, and carry the procedure into a cron run, a terminal session, or a message from Slack. Profiles isolate separate agents. Toolsets, plugins, MCP servers, webhooks, and provider switching give the runtime a broad execution surface. The current documentation describes the product as an autonomous agent that gets more capable the longer it runs. That claim maps to an architectural center, not a paragraph on a pricing page.</p><p>Compounding changes the economics when the work repeats with variation. A deployment workflow is never identical twice, but the repository conventions, verification commands, rollback rules, and known failure modes accumulate. Research jobs change topics, while the source-quality rules and filing process persist. Operational agents encounter the same systems through different incidents. A runtime that can preserve those procedures stops spending tokens rediscovering the environment and starts spending them on the part that changed.</p><p>The cost arrives in governance. A skill is executable institutional memory. A bad skill can repeat a bad decision with more confidence each time. A stale memory can survive longer than the system it describes. Provider choice does not remove the need for model evaluation. Gateway reach does not remove the need for channel-specific authorization. Teams adopting Hermes need skill review, memory provenance, profile boundaries, tool allowlists, and a promotion path from learned procedure to trusted procedure. Self-improvement without review is drift with a flattering name.</p><p>Hermes also imposes a runtime boundary. It is an environment you operate, not a small orchestration library hidden inside an existing service. That is an advantage when the agent itself is the long-lived worker. It is friction when the agent is one bounded component inside a product with established deployment, identity, and audit systems. The question is not whether Hermes can call the API. It can. The question is whether you want Hermes or your application to own the loop.</p><p>The strongest architecture is often the option missing from the three-way debate: both.</p><p>Keep the regulated or revenue-bearing workflow in a service you own. Give it typed inputs, idempotent operations, explicit approval states, and an audit trail. Use LangGraph or CrewAI Flows inside that service if their graph abstractions earn their place. Put a runtime at the edge where unstructured work begins. OpenClaw can normalize channels and route an event into the service. Hermes can research, prepare the packet, invoke the service through a narrow tool, learn the surrounding procedure, and deliver the result. The runtime handles context and continuity. The service handles authority.</p><p>That split prevents the worst failure in each direction. The custom runtime does not have to rebuild messaging, memory, scheduling, and tool discovery. The adopted runtime does not become the system of record for a decision it cannot deterministically replay. You can replace the edge runtime without rewriting the domain transaction, and you can change the workflow without erasing the agent&#8217;s accumulated operating knowledge. Flexibility comes from a clean ownership line, not from avoiding dependencies.</p><p>The decision rule is narrower than most scorecards make it. Build when the execution topology contains your business rules and you are prepared to operate every failure around it. Adopt OpenClaw when reach, TypeScript integration, and shared channel maintenance matter more than accumulated procedure. Bet on Hermes when the agent&#8217;s value should compound across sessions, tools, and recurring work, and you are willing to govern what it learns. Combine them when the agent needs broad judgment but the business decision needs a deterministic owner.</p><p>Tomorrow&#8217;s scorecard will compare the field across persistence, recovery, memory, tools, community, token efficiency, and cost. Those dimensions matter. None of them rescues an architecture that assigned authority to the wrong layer.</p><p>Your runtime choice is the failure boundary you agree to own at 2:13 a.m.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Week in Agent Runtimes That Actually Mattered]]></title><description><![CDATA[The protocol named agent identity.]]></description><link>https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-230</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-230</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Fri, 28 Aug 2026 10:17:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!g5j5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The protocol named agent identity. Memory washed the flags off an email. Three MCP servers listened on every interface. Nvidia reportedly bought the warehouse.</em></p><div><hr></div><p>The tool middleware week did not fail at discovery. It failed at who the caller is, what the memory file believes, and which interface the server bound.</p><p>On Saturday the Model Context Protocol maintainers published the updated roadmap. David Soria Parra and Den Delimarsky, six minutes of prose, five priority areas. The one that matters for this publication is the third: agent identity and enterprise-ready security. MCP authorization today is built around a person approving access in a browser. That works for a human with a tab open. It does not work for an agent running as a cloud workload, acting for a user who is not present, or handing a narrower grant to a sub-agent.</p><p>The maintainers named the path. Finish Demonstrating Proof of Possession. Drive Workload Identity Federation. Use the ID-JAG grant behind Enterprise-Managed Authorization. Standardize token exchange. They will keep showing up in the IETF OAuth working groups so those building blocks exist in the standards, not only in a blog post. Tuesday&#8217;s piece on this desk said the catalog finds the tool and the hook decides whether it fires, and neither one can tell you who the agent is when the call leaves the process. The roadmap is the first time the USB-C of agent capabilities has admitted that sentence is a spec gap, not a vendor feature request.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g5j5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g5j5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g5j5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1500122,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/213123186?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!g5j5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!g5j5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc922cca2-cbca-4fc5-9f83-715a199c0ca0_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The rest of the document is the operational half of the same problem. The 2026-07-28 specification already dropped protocol-level sessions and the initialization handshake so a server can scale without holding state. Clients can call server/discover before they do anything else. List results are cacheable. Tasks moved into an official extension. A new Multi Round-Trip Requests pattern replaced server-initiated requests so elicitation works on a stateless server. Next they want Streamable HTTP covering local servers over stdio, one transport instead of a zoo. They also want progressive discovery, because connecting to a server with a hundred tools means the model pays for the entire surface before the user has asked a question. HN put 270 points and 161 comments on the post. The split was the expected one: HTTP-native cheers, sampling mourners, people who already lazy-load and are moving to code-mode. None of that changes the identity line.</p><p>The identity line is also where the week&#8217;s most useful attack landed, and it did not need a CVE.</p><p>Patrick Walsh at IronCore Labs spent April sending prompt injections at an OpenClaw instance on GPT-5.4. Direct ones failed. OpenAI&#8217;s March detector claimed 99.8 percent. OpenClaw wrapped untrusted email in &lt;&lt;&lt;EXTERNAL_UNTRUSTED_CONTENT&gt;&gt;&gt; tags, stripped angle brackets, and taught the model to ignore instructions inside the wrapper. The hourly summarizer named the attempts: coordinated prompt-injection, a cute little test wrapped in a poem, leave it for the user. Spotlighting worked on the ingest path.</p><p>It did not survive promotion. Daily memory files ingested excerpts of those emails without the untrusted tags, without the trust level, without the source. Dreaming, OpenClaw&#8217;s periodic distillation of daily notes into MEMORY.md and USER.md, is the second wash. Those two files load into every context, including HEARTBEAT and cron, the sessions that never saw the original email. Walsh sent a series of messages declaring the attacker address an internal account, the user&#8217;s other address, a bug in the external list. After the third message he did not have to wait for dreaming. The summarizer cron updated the current day&#8217;s memory on its own. Then commands from that address were instructions from an internal alt, not a standard email. The agent tried to send mail, failed, and followed the CLI directions.</p><p>The detector was not the hole. The memory writer was. A 99.8 percent injection score on ingest is theater if promotion strips the flag. Walsh names ChatGPT, Claude, and Hermes in the lede for the same pattern: spotlight or flag the untrusted text, write a memory file, drop provenance on the way in. Tuesday said the missing piece is a standard for agent-level OAuth that does not require a human to click Allow every 90 days. Thursday&#8217;s OpenConnector piece said the enterprise question is not whether the agent can find the tool, but whether it should be allowed, and whether you can prove it. Walsh&#8217;s writeup is the memory version of that question. Diff MEMORY.md the way you diff IAM. Cron should not load user-authored memory that started life in an untrusted channel.</p><p>The implementations of the connector spent Thursday proving the other half.</p><p>NVD assigned a 17:20Z wave against MCP HTTP servers that defaulted to every interface with auth off. UI-TARS-desktop, CVE-2026-81735, scored 10.0. @agent-infra/mcp-http-server listened on :: when no host was given. Auth middleware applied only if the caller supplied it. The commands server handed run_command to child_process.exec on the caller&#8217;s string. Anyone who could reach the port ran commands as the desktop user. The listen default moved to 127.0.0.1 in commit c2ad42e. The package version stayed 1.2.4 across that change. The pin is the commit, not a tag. mcp-router, CVE-2026-81094, scored 9.3: serve bound the aggregator to all interfaces on a fixed port and required a token only when the flag was passed. Release 0.6.3 defaults to loopback and refuses to start without a token when the host is not loopback. Telnyx MCP, CVE-2026-81098, scored 9.3 through 6.83.0. Current npm telnyx is 7.17.0, loopback, required server API key. Same costume we have been writing down all month. USB-C with the port open and the lock optional is not a connector. It is a LAN shell.</p><p>If the protocol is going to be the default answer for tool discovery, the default listen address is part of the product. Stdio until you have a token and a loopback bind. Inventory 0.0.0.0 and :: on every MCP HTTP process you actually run.</p><p>The warehouse underneath those tools changed owners, if the report holds.</p><p>The Information, via CNBC and TechCrunch, said Nvidia agreed to buy Hugging Face for $12.9 billion. Neither company has commented. Business Insider still will not call the ink dry. Treat the number as a reported agreement, not an 8-K. What they would be buying is the two-sided Hub, the default from_pretrained libraries, the telemetry of who downloads which shard onto which GPU, and a cloud re-entry after last year&#8217;s DGX Cloud pullback. <strong><a href="http://ggml.ai/">ggml.ai</a></strong> joined Hugging Face around February. Local GGUF already sits under this roof. Hugging Face raised at $4.5 billion in 2023, turned down a $500 million Nvidia check at $7 billion late last year because it did not want a dominant investor, and was recently doing about $150 million a year. Stripe took OpenRouter earlier this month. Same pattern, one week apart: the meter and the warehouse get bought, the models stay &#8220;open.&#8221;</p><p>Do not rewrite CI tonight. The actionable move is the one that was already hygiene. Mirror the checkpoints you actually serve so Hub origin is a convenience, not a single point of failure. Watch GGUF, MLX, and non-CUDA defaults, not a fantasy TOS that bans quants. Pin the llama.cpp commit you ship, independent of Hub&#8217;s default branch. Model IDs will keep resolving. The operator behind the catalog is the thing that moved.</p><p>The quiet open-source drop sat under all of that. Munder Difflin is an Electron desktop that wraps the terminal CLIs you already pay for (Claude Code, Codex, Gemini CLI, Qwen, OpenCode, Copilot, Cursor, local models) and seats each one as a real process on an office floor. You talk to one clone. That clone routes. Everyone else is a worker with a desk and a mailbox. The coordination layer is a local git repo of plain files. Agents write outbox/. A harness router delivers into inbox/. No agent touches git. Single-committer, on purpose, because twelve CLIs trying to commit at once corrupt index.lock. GitHub shows 5,102 stars this morning. Tagged v0.4.6 on August 27. v0.4.5 was the release that admitted the load-bearing bugs: cost reporting reset on every restart, semantic memory on Apple Silicon returned NaN embeddings, mail sat in inboxes nobody woke. Official builds send anonymous usage events. Opt out, or build from source.</p><p>Treat it as a CLI multiplexer with a file-shaped mailbox, not as The Office for agents. If you already pay for two of those CLIs, the new cell is the hive directory, which you can read because it is git. Last week&#8217;s OneCLI put the credential outside the model. Munder put the conversation between tools into a repo you own. Both are votes for the same design: the middleware is files and a gateway, not a prompt.</p><p>The pattern across the five is the split this week spent drawing. Discovery is no longer the scarce layer. Identity is. MCP named the grants. Walsh showed what happens when memory promotion drops the grant on the floor. The CVE wave showed what happens when the HTTP transport defaults to every interface. Nvidia buying the Hub, if it closes, is who hosts the weights those tools load. Munder is the local reminder that the tools were already on the machine.</p><p>Next week the arc moves from the connector to the decision that sits on top of it. Build your own runtime, adopt OpenClaw, or bet on Hermes. The scorecard after that. The week after the USB-C thesis is the week you pick which socket you are willing to live with.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[OpenConnector]]></title><description><![CDATA[The Tool Integration Layer Built for Enterprise API Governance]]></description><link>https://signalovernoise.tech/p/openconnector</link><guid isPermaLink="false">https://signalovernoise.tech/p/openconnector</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Thu, 27 Aug 2026 10:15:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zJlR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The catalog finds the tool. The gateway decides whether this token may call it, and writes down why.</em></p><div><hr></div><p>The catalog answers whether the agent can find the tool. OpenConnector answers whether this token is allowed to call it, and whether you can hand an examiner the record afterward.</p><p>Yesterday I closed the middleware stack on a rotation test. If the secret lives in the process, a rotation is an outage. The missing piece on that diagram was not another registry. It was a gateway that keeps the credential, applies a policy, and writes a log you can verify. That is the job OOMOL Lab shipped as OpenConnector.</p><p>If you still have github.com/openconnector/openconnector in a bookmark, throw it out. That repo does not exist. The tree is oomol-lab/open-connector, Apache 2.0, TypeScript, Node 22 or newer. It was created June 29. This morning it sits at 5,337 stars and 454 forks, last push about an hour ago. The latest tag is v1.4.0, published August 20. Forty-six commits have landed on main since that tag. The live catalog at connector.oomol.com reports 1,445 providers and 14,791 actions. Two months ago the research note in this vault called the project twenty days old and sitting near 3,000 stars. The count moved. The architecture did not: credentials stay behind the runtime boundary, and agents get schemas and results.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zJlR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zJlR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zJlR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115437,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212975014?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zJlR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zJlR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6eefbdd6-a2da-4460-970d-406ef8eec1a7_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The packaging is the first trap. The marketing site and the open-source runtime are not the same product with different wrapping.</p><p>openconnector.dev is a private-beta waitlist that pitches a Composio alternative with a one-line base URL change and an SDK named @open-connector/sdk. That package 404s on npm this morning. The client that actually installs is @oomol-lab/connector 1.1.0, published August 6, MIT, 21 stars, in oomol-lab/connector-sdk. The CLI lives in a third repo, oomol-lab/oo-cli, at v1.7.10 as of this morning. If you are evaluating the governance claim, read the self-hosted runtime. The landing page is a waitlist.</p><p>The runtime is a gateway. You point an agent at MCP, HTTP, or OpenAPI. The agent never sees a GitHub PAT or a Gmail refresh token. It presents a runtime token, names an action, optionally names a connection alias, and gets a result. The gateway decrypts the stored credential with AES-256-GCM, injects it server-side, and returns the upstream response. Admin surfaces take OOMOL_CONNECT_ADMIN_TOKEN. Agents take an oct_&#8230; runtime token, shown once, stored as a hash. Mix those up and the console works while /v1 and /mcp return unauthorized, or the reverse. The docs split those tokens on purpose. Treat a 401 on the wrong surface as a configuration error, not an SDK bug.</p><p>MCP is Saturday&#8217;s USB-C connector used as a discovery surface rather than a dump of fourteen thousand tools into the prompt. The local server at /mcp speaks the 2026-07-28 protocol as of v1.4.0. Stateless JSON-RPC POST. No hanging SSE stream. Five tools:</p><ul><li><p>list_apps</p></li><li><p>list_connections</p></li><li><p>search_actions</p></li><li><p>get_action_guide</p></li><li><p>execute_action</p></li></ul><p>That is the Composio meta-tool idea executed on a process you run. Search, fetch the guide, then execute. The guide is markdown: input schema, required scopes, the safe account label for the connection you selected. Omitting connectionName uses default. A named connection that does not exist does not silently fall back. Persistent tokens with a non-empty allowedConnections list never see the other accounts in discovery. Denied calls fail with HTTP 403 or MCP connection_not_allowed before credential lookup. That last sentence is the product.</p><p>Policy has two layers, and you want both.</p><p>Deployment policy is environment. OOMOL_CONNECT_ALLOWED_ACTIONS and OOMOL_CONNECT_BLOCKED_ACTIONS. Their own docs use github.* with github.delete_repository on the block list. I would ship that on day one. Blocked wins.</p><pre><code><code>OOMOL_CONNECT_ALLOWED_ACTIONS="github.*"  \
OOMOL_CONNECT_BLOCKED_ACTIONS="github.delete_repository"  \
docker  compose  up  --build</code></code></pre><p>Runtime tokens add a second cut: allowedActions, blockedActions, allowedProxies, allowedConnections. Proxies default to empty, so a token cannot call /v1/proxy/:service until you grant a provider or *. Bootstrap tokens and JWTs have no stored grant. They inherit only the deployment policy. Stand this up on a public origin with the bootstrap token and a wide allow list, and you have rebuilt the PAT-in-the-environment problem with extra YAML.</p><p>v1.4.0 is the release that makes the OAuth side less embarrassing. Connection-scoped OAuth apps. Clients can request a scope subset. Gmail dropped the admin-only scope from user OAuth. Figma went to public scopes plus PKCE. Run-log redaction got three passes in the same tag because URL-shaped secrets kept leaking into summaries. That is the work a gateway has to do, and it is the work a hosted catalog can hide until an incident.</p><p>The audit trail is the claim that will get a security team in the room. Every credential use writes an append-only NDJSON record under .audit/, hash-chained with SHA-256 over the canonical content plus the previous hash. Actor, target, action, outcome. No token values, no request bodies. A verifier in the repo re-derives the chain. Set OTLP_ENDPOINT and Loki becomes the queryable mirror. The on-disk file is the integrity artifact.</p><p>Here is where I would not let a vendor slide. The default hash-chain state is in-memory and single-process. Restart the process without the shared-state strategy they document, and the chain does not mean what you told the auditor it means. There is no audit read or export HTTP API. Long-term tenant-scoped export is gated on an @open-connector/ee-audit license that is a follow-up, not shipping. Encryption-key rotation has no command. Change CONNECTOR_ENCRYPTION_KEY without re-encrypting and every stored credential is gone. Phoenix traces a span. This journal is closer to what I asked for on Tuesday. Closer is not an examiner signing off.</p><p>The identity hole from Tuesday is still a hole. Agents authenticate with a project-scoped API key or a runtime token. That attributes a call to a project. It does not mint an agent principal that can acquire a short-lived, audience-bound token without a human in the loop. First connection for OAuth is still a browser. The user clicks Allow. The refresh token lives in your vault instead of Composio&#8217;s. That is a real improvement for custody. It is not agent-scale authentication. Fill the blank with a shared service account behind this gateway and you will still fail the CISO question after the first bad tool call. You will fail it with better logs.</p><p>Composio remains the catalog I would rent for the long tail I refuse to operate. 29,879 stars yesterday, hosted OAuth, seven meta-tools, a session tied to your user_id. You do not run Postgres. You also do not own the credential, the policy, or the log. Strands remains the in-process hook for the call that is the product: BeforeToolCallEvent on a refund you wrote. OpenConnector sits between those two on purpose. Self-host it when the SaaS tail has to live in your VPC, when allow/block has to be independent of the model&#8217;s tool list, and when a hash-chained record of who called what is a requirement rather than a nice-to-have. Skip it when you have five MCP servers you already wrote. Skip it when the integration is load-bearing internal code and the gate belongs in your process. Skip the hosted waitlist until the SDK on the homepage exists on npm.</p><p>I would stand it up as Docker Compose on a private network, encryption key in a secrets manager, ALLOWED_ACTIONS set to the three providers the agent is allowed to see, BLOCKED_ACTIONS set to every delete and every send, one runtime token per agent with allowedConnections pointing at the work GitHub account, MCP pointed at /mcp. Then rotate the GitHub PAT in the console and confirm the next github.get_current_user still works without bouncing the agent.</p><pre><code><code>curl  -s  -X  POST  http://localhost:3000/v1/actions/github.get_current_user  \
-H 'authorization: Bearer RUNTIME_TOKEN' \
-H  'x-oo-connector-alias: work'  \
-H 'content-type: application/json' \
-d  '{"input":{}}'</code></code></pre><p>That is yesterday&#8217;s test, applied to this gateway. If that call dies with the old key, you bought a catalog with extra ceremony.</p><p>The connector is settled. The catalog is still a rental unless you run it. The hook is still yours. The gateway is how you say no, and how you prove you said no after the model has already tried.</p><p>Friday I will pick the runtime news that actually moved this week. The decision tonight is narrower than whether OpenConnector is &#8220;the enterprise Composio.&#8221; Can you name the agent, the user, the action, and the policy that allowed it, without opening the process?</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Tool Middleware Stack]]></title><description><![CDATA[A Reference Architecture for Agent-Tool Integration]]></description><link>https://signalovernoise.tech/p/the-tool-middleware-stack</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-tool-middleware-stack</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Wed, 26 Aug 2026 10:03:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dJ1a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>If the secret lives in the process, a rotation is an outage.</em></p><div><hr></div><p>The secret that sits in the agent&#8217;s environment is the outage you have not had yet. Rotate it and the process dies, or worse, keeps running on the old key until the vendor rejects it in the incident channel. Put the same secret behind a proxy the agent never sees, and a rotation is a vault write. The next call picks up the new material. The loop never notices.</p><p>That is the design decision this stack is built around. Everything else this week was the wiring.</p><p>This week we named the pieces. MCP is the connector. Composio is the catalog that keeps a thousand wrappers out of the prompt. Strands is the in-process hook that can cancel a refund after lookup and before the socket opens. Yesterday I split authentication into three problems teams keep collapsing: client to server on the MCP transport, credential brokering so a prompt-injected agent cannot walk out with a PAT, and an agent principal that can acquire a short-lived token without a human clicking Allow every ninety days. The first two have products. The third does not. This architecture leaves that blank named, and builds the rest so a rotation cannot take the agent down while it is still blank.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dJ1a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dJ1a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dJ1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:93687,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212825439?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dJ1a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!dJ1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd208135b-f89c-4b55-8dab-153e8587fe93_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The requirement that defines this class is not &#8220;the agent can call tools.&#8221; Any tutorial can do that. The requirement is that the agent can discover a tool it did not ship with, invoke it without holding the credential, survive the credential changing under it, and leave a trace an engineer can replay. If you also need a record a regulator will accept, you are still in the July 27 gap. Phoenix and Langfuse will not close it. I am not going to smuggle an audit trail into a middleware diagram and call it solved.</p><p>Discovery is the layer that is, for the first time, a solved class of problem. The Model Context Protocol is the port. The official Python SDK is at 2.1.1, released August 25. The community registry sits at 7,191, last push August 22. A server advertises a name, a description, and an input schema. A client discovers it. After the July 28 spec, a single HTTP request can call a single tool with no session to babysit. That is the connector. It is not the catalog.</p><p>The catalog is what keeps discovery from eating the context window. ComposioHQ/composio is at 29,879 stars this morning, MIT. The Python package is composio 0.20.0, uploaded August 18. TypeScript is @composio/core 0.17.0, same day. You create a session scoped to a user_id from your own database. The model gets seven meta-tools. It searches, fetches schemas for the slugs it needs, and executes. It does not receive a thousand function definitions on turn one. If you only have five first-party tools, you do not need this. Write five MCP servers and register them in the runtime. The catalog earns its keep on the long tail of SaaS: Gmail, Jira, the payer portal, the CRM you will not wrap yourself. The runtime calls tools. It does not contain them.</p><p>Authentication is where the outage lives. Treat it as three layers because collapsing them is how the incident starts.</p><p>Client-to-server belongs on the MCP transport. OAuth 2.1, Protected Resource Metadata, resource indicators, short-lived tokens. Refuse servers that skip it. Strands 1.53.0 shipped client OAuth on streamable HTTP on August 21. The harness-sdk monorepo is at 7,009 stars, with Python strands-agents 1.53.0 and TypeScript @strands-agents/sdk 1.14.0. That authenticates the MCP client to the MCP server. It does not authenticate the agent to GitHub.</p><p>Credential brokering belongs under the agent. Infisical Agent Vault is at 2,137 stars, v0.39.1 on August 4, last push August 23. It sits on the wire as a TLS-intercepting proxy, swaps a placeholder for the real key, and never lets the secret enter the agent&#8217;s context. You own the CA. You own the egress firewall. Harbor SDK is the other cut: the agent invokes an abstract tool, the connector holds the key in another process, and the agent never learns the endpoint as a network primitive. Harbor is at 384 stars. The latest tag is still v0.1.2, published June 5. Nothing has shipped since. Treat it as a pattern. Implement the pattern in a process you control, or use Agent Vault if you can live with TLS interception.</p><p>The rotation rule sits here, and it is the one I would not negotiate. The agent never holds the secret. Not in an environment variable. Not in a config file it can read. Not in a vault it can query, because a query puts the value in context, and context is the prompt-injection surface. The agent sends a request with a handle. The proxy or the connector substitutes at call time. A human rotates the key in the vault. The next outbound call uses the new material. The process does not restart. The session does not die. If your architecture requires a bounce to pick up a new PAT, you do not have agent-scale authentication. You have a cron job with extra steps.</p><p>Agent identity is the blank. Entra Agent ID covers Copilot-shaped estates. Hermes on a desktop, a Strands loop on Lambda, and a Composio session for the SaaS tail do not share a principal. Until they do, bind user-delegated SaaS to a real user_id, put a rotation alarm on the refresh token, and keep the call that is the product on a path you can refuse.</p><p>Execution is the hop after the hook. Three shapes, pick by blast radius. A local MCP server is a subprocess with the permissions of whatever started it. That is fine for a desktop agent reading your notes. It is not fine for a refund tool pointed at production. AWS Lambda, or an equivalent isolated function, is the shape I want for anything irreversible: one credential, one role, one timeout, no filesystem. A direct API call from inside the loop is the fastest path and the one I would reserve for reads you can afford to replay. Strands&#8217; BeforeToolCallEvent is the gate on the in-process path. Evaluate the intersection, not the union: what this agent is allowed to do, and what this user is allowed to do, per action. Composio&#8217;s hosted MCP path skips the SDK hooks. Do not put the refund on that path. Search and schema fetch can live there. The write that moves money stays in-process, behind the cancel.</p><p>Observability is the last layer, and it is debug, not evidence. Arize Phoenix 20.4.0 shipped August 26. It traces agents as trees: model call, tool call, retry, the span that shows which hook cancelled the refund. Langfuse&#8217;s Python SDK is 4.14.5, uploaded August 24. The platform tagged v4.19.0 on August 25. Phoenix for depth, self-hosted, inside the VPC. Langfuse when you want prompt versions, a simpler deploy, or an EU-hosted option. Instrument both if you have the audience split. Instrument Phoenix alone if you are still the only person who will open the trace. Neither one is an audit trail. I wrote that in July. The examiner wants which agent, acting for which user, called which tool, against which resource, under which policy, and why the runtime allowed it. A span is not that record. Keep the hashed append-only log if you are in the regulated case. Keep Phoenix so you can fix the loop.</p><p>Where this flexes is the parts you can swap without breaking the class. Skip Composio until the fifth SaaS toolkit. Five MCP servers you wrote are cheaper than a hosted catalog you do not yet need. Skip Agent Vault if every tool is a connector you own and the secret never enters the process; that is the Harbor pattern without the dormant SDK. Skip Lambda if the agent is personal and the worst case is a bad file in a sandboxed directory. Skip Langfuse until someone asks for a prompt version. The fixed parts do not flex. Discovery outside the runtime. Secrets outside the process. A hook that can refuse. Traces you can replay. Drop the secret into the agent&#8217;s environment and a rotation is an outage, no matter how pretty the registry looks.</p><p>What it costs is mostly operational discipline, then rent, then a CA. Every new tool needs a home in one of the three execution shapes before it ships. Every Composio connection needs a user_id that will still mean something when the engineer who clicked Allow leaves. Agent Vault needs certificate distribution and an egress deny that makes HTTPS_PROXY a boundary instead of a suggestion. Phoenix needs a Postgres and someone who will look at a trace at 2 a.m. Harbor-as-a-pattern needs a process boundary you actually maintain. I would not stand all four layers up on a Friday against a live billing API. I would stand discovery and the hook against a read-only tool, then move the secret out of the environment, then add traces, then let the first write through.</p><p>The starting point is smaller than the diagram. One MCP server you own. One runtime hook that can cancel. One secret in a vault or a connector, never in the process. One Phoenix instance, even if it is a container on the same box. Add Composio when the long tail appears. Add Agent Vault when the agent has a filesystem and you cannot trust the next document. Add Langfuse when prompt versions become an argument. Leave the agent-principal blank on the diagram on purpose. If you fill it with a shared service account, you will not be able to answer the CISO after the first bad tool call.</p><p>Thursday I will look at OpenConnector, which treats the question as governance rather than catalog: not can the agent find the tool, but should this agent be allowed to use it, and can you prove that to an auditor. Tonight the test is plainer. Rotate a key. If the agent keeps running, you built the stack. If you have to bounce the process, you built a demo with a calendar reminder.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Tool Problem Nobody Talks About]]></title><description><![CDATA[Authentication at Agent Scale]]></description><link>https://signalovernoise.tech/p/the-tool-problem-nobody-talks-about</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-tool-problem-nobody-talks-about</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Tue, 25 Aug 2026 10:20:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oM1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>MCP will find the tool. It will not vouch for the hand that calls it.</em></p><div><hr></div><p>The catalog finds the tool. The hook decides whether it fires. Neither one can tell you who the agent is when the call leaves the process.</p><p>That is the hole this week has been circling. MCP made tools discoverable. Composio put a thousand of them behind seven meta-tools and a session. Strands put a cancel hook on the path you already own. Discovery is, for the first time, a solved class of problem. Authentication is not. Every one of those tools still opens with a credential minted for a human, stuffed into an environment variable, or refreshed by a browser tab someone has to keep clicking.</p><p>I keep watching teams collapse three different auth problems into one sentence, then wonder why the sentence does not hold.</p><p>The first problem is client to server. The MCP 2026-07-28 spec finally treats an MCP server as an OAuth 2.1 resource server. Protected Resource Metadata is required. Resource Indicators bind a token to one audience, so a hostile server cannot harvest a token meant for another. Client ID Metadata Documents replace the Dynamic Client Registration mess. Strands 1.53.0 shipped client OAuth on streamable HTTP on August 21. That is real work. It authenticates the MCP client to the MCP server. It does not authenticate the agent to GitHub, or to the payer portal, or to the billing API the refund tool is about to hit.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oM1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oM1G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oM1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1352324,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212678649?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oM1G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!oM1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84679161-c791-4781-adf7-46d34b9a64ef_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The second problem is credential exfiltration. We covered that in July. Infisical Agent Vault sits on the wire as a TLS-intercepting proxy, swaps a placeholder for the real key, and never lets the secret enter the agent&#8217;s context. The repo sits at 2,129 stars, up from about 1,900 a month ago. v0.39.1 landed August 4. Last push was August 23. Harbor SDK takes the other cut: the agent invokes an abstract tool, the connector holds the key in another process, and the agent never learns the endpoint exists as a network primitive. Harbor is at 385 stars. The latest tag is v0.1.2, published June 5. Nothing has shipped since. A proxy that hides the key and a sandbox that hides the tool both answer &#8220;how do I stop a prompt-injected agent from walking out with a PAT.&#8221; They do not answer &#8220;how does this agent prove it is allowed to act.&#8221;</p><p>The third problem is the one nobody has a protocol for. The agent itself needs an identity. Not a service account copied from the CI pipeline. Not the intern&#8217;s Google login that Composio asked them to approve in a browser. A principal the authorization server can issue a short-lived, audience-bound token to, without a human in the loop every ninety days.</p><p>Composio&#8217;s first-connection flow is the honest picture of where the industry sits. The agent needs Gmail. Composio generates an OAuth link. A human opens a browser, clicks Allow, and the connection sticks to that user_id. After that, refresh tokens keep it alive until they don&#8217;t. When Google or Microsoft forces reconsent, or the refresh token is revoked, or the engineer who clicked Allow leaves the company, the agent stops. That is a solved onboarding ceremony for a product with users. It is not agent-scale authentication. It is a human workflow with an SDK wrapped around it.</p><p>The reason this keeps getting papered over is that OAuth was designed for a person sitting in front of a browser. PKCE protects the code exchange. It does not prove which workload asked. If any process can start a PKCE flow, the flow is not authenticating an agent. It is authenticating whoever showed up. Infrastructure-asserted identity is how we already solve this for services: an IAM role on a Lambda, a Kubernetes service account token projected into a pod, a federated credential on a workload identity. The authorization server attests the runtime, then mints an ephemeral token. No long-lived secret. No ninety-day click. Agents do not get that treatment because we keep modeling them as chat users who happen to call APIs.</p><p>Microsoft noticed. Entra Agent ID went generally available in April. Agent identities in Entra do not hold credentials of their own. They acquire tokens through an identity blueprint using federated credentials. That is the right shape: the agent is a principal, the secret lives elsewhere, Conditional Access can attach to it. It is also a Microsoft-shaped answer. If your estate is Copilot and Agent 365, you have a path. If your estate is Hermes on a desktop, a Strands loop on Lambda, and a Composio session for the SaaS tail, Entra does not cover the call that will page you.</p><p>The audit problem sits under all three. An examiner does not want to know that a token was valid. They want to know which agent, acting for which user, called which tool, against which resource, under which policy, and why the runtime allowed it. Agent Vault logs destination, status, latency, and actor name. That is a request log, which is better than nothing and not an authorization decision. Strands can cancel a refund in BeforeToolCallEvent and leave a trace. Composio&#8217;s hosted MCP path skips the SDK hooks, so the reason code never gets written. MCP itself still treats authorization as optional. A server that never asks for a token is spec-compliant. Spec-compliant and unauthenticated is how you get a public MCP port with a production database behind it.</p><p>Here is what I would actually build while the standard catches up.</p><p>Treat the three problems as three layers, because collapsing them is how the incident starts. Client-to-server auth belongs on the MCP transport: OAuth 2.1, resource indicators, short-lived tokens, refuse servers that skip it. Credential brokering belongs under the agent. Use Agent Vault if you want a network guarantee and you can own the CA and the egress firewall. Treat Harbor&#8217;s connector model as a pattern if the blast radius of a free-form tool call is the thing you cannot accept, and treat the SDK as a pattern too: eighty days without a tag is not a product you bet a control plane on. Agent identity belongs in the runtime you already run. Internals should use the workload identity the platform already knows how to attest. User-delegated SaaS can keep the browser consent if you bind it to a real user_id and put a rotation alarm on the refresh token instead of discovering expiry in the incident channel. The call that is the product stays in-process, with the intersection in the hook: what this agent is allowed to do <em>and</em> what this user is allowed to do, evaluated per action. Never the union. If a user cannot issue a refund, the agent acting for them cannot either, no matter what IAM role the process was launched with.</p><p>The decision rule is simpler than the architecture. If a human has to click Allow for the agent to keep working, you do not have agent-scale authentication. You have a demo with a calendar reminder. If the agent holds a static key, you have a prompt-injection exfil waiting on the next poisoned document. If the only identity in the audit log is a shared service account, you cannot answer the question a CISO will ask after the first bad tool call.</p><p>Tomorrow I will put the layers on one diagram: discovery, authentication, execution, observability, and the one design choice that decides whether a credential rotation takes the agent down. Leave a blank where the agent&#8217;s name should be and you will feel the outage before you feel the architecture.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Strands Agents SDK]]></title><description><![CDATA[AWS&#8217;s Quiet Bet on the Tool Middleware Pattern]]></description><link>https://signalovernoise.tech/p/strands-agents-sdk</link><guid isPermaLink="false">https://signalovernoise.tech/p/strands-agents-sdk</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Mon, 24 Aug 2026 09:39:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!22x-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The catalog finds the tool. The hook decides whether it fires.</em></p><div><hr></div><p>I called Strands a framework in May. It still is. The reason it belongs in a week about tool middleware is the hook that sits between the model&#8217;s intent and the call.</p><p>Yesterday&#8217;s post was Composio: a hosted catalog, a session scoped to one of your users, seven meta-tools that keep a thousand wrappers out of the prompt. MCP is the connector those wrappers travel over. Strands is the other shape of the same problem. You own the loop. You decorate the functions. You register a callback that can cancel a refund before the SDK ever opens a socket. Nobody is pretending this is npm. It is the in-process gate.</p><p>The numbers moved since May, and the repo name moved with them. In May the public tree was strands-agents/sdk-python, version 1.38.0, about 5,800 stars. That URL now 301s to strands-agents/harness-sdk. The monorepo holds the Python SDK, the TypeScript SDK, the docs site, and a bundled MCP server. This morning it sits at 6,990 stars, Apache 2.0, last push a few hours ago. The Python package on PyPI is strands-agents 1.53.0, uploaded August 21. The TypeScript package is @strands-agents/sdk 1.14.0, same day. Cadence since late July has been roughly weekly: 1.50.0 on July 24, 1.51.0 on August 7, 1.52.0 on August 12, 1.53.0 last Friday. That is not a science-fair SDK. It is a product with customers inside Amazon and a public tree that has to keep up.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!22x-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!22x-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!22x-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!22x-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!22x-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!22x-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/357b2276-acf6-4761-860c-64119689ef57_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1878271,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212522280?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!22x-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!22x-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!22x-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!22x-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F357b2276-acf6-4761-860c-64119689ef57_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If a search result still points at github.com/strands-ai/strands-agents-sdk, that repo never existed. You want the monorepo.</p><p>I am not going to rerun the May teardown. The model-agnostic claim still holds. I have moved a Strands agent from Bedrock to the Anthropic API without touching the loop or the tool definitions, and that is still the test. OpenTelemetry spans still come out of the box for model calls, tool invocations, and loop iterations. Bedrock is still the default provider. Lambda, Fargate, EKS, and AgentCore still lead the deploy docs. The lock-in is still gravitational, not contractual. You can run this on a laptop with Ollama and never open the AWS console. Most teams that pick it will not.</p><p>What changed is where the project spends its engineering. The community tools package, strands-agents-tools 0.8.6, shipped August 7 with a deprecation pass and a warning in the README that the tools are experimental and need an independent security review before production. Sleep, editor, and shell now point at SDK-vended replacements. Batch is gone because concurrent execution is the default. Think, current_time, memory, and retrieve point at native reasoning config, a context injector, or MemoryManager. The maintainers say the repo will eventually be archived. They would rather you use an official vendor MCP server than a second wrapper they have to keep in sync. That is the tell. The interesting surface is no longer a pile of community tools. It is the SDK&#8217;s own tool contract, plus MCP as a first-class source, plus the hook that can refuse a call.</p><p>A tool in Strands is a typed function. In Python you decorate it. In TypeScript you pass a Zod schema. The model sees a name, a description, and an input schema. Drop a .py file on a path and the Python SDK will load it, which is still the fastest inner loop I have used in this category and still a footgun if the file has hidden state. The model asks for a tool. The executor validates the request, runs the function, and feeds the result back. Failures come back as tool errors, not as exceptions that kill the loop. That part is ordinary. Every serious SDK does a version of it.</p><p>The middleware is the event that fires after lookup and before execution.</p><pre><code><code>from strands import Agent, tool
from strands.hooks import BeforeToolCallEvent  

@tool
def  issue_refund(order_id: str, amount: float) -&gt; str:
    """Process a customer refund against the billing record."""

    return payments.refund(order_id, amount)  

def  approve_refunds(event: BeforeToolCallEvent):
    name = event.tool_use["name"]

    if name == "issue_refund":
        event.cancel_tool = "Refunds require a human."  

agent = Agent(tools=[issue_refund], hooks=[approve_refunds])

agent("Refund order 1842 for $86.40")</code></code></pre><p>BeforeToolCallEvent can inspect the call, rewrite it, cancel it, or raise an interrupt and hand control back to a human. BeforeToolsEvent does the same for the whole batch. On resume, completed results are kept so the model is not asked the same question twice. TypeScript names the cancel field event.cancel. Python names it event.cancel_tool. Same gate. The homepage sample uses it to refuse a report that has no citations. I would use it to refuse a write against production unless a ticket id is present in the arguments. That is tool middleware. It is not a catalog. It is policy sitting on the call path you already own.</p><p>MCP plugs into the same path. You construct an MCPClient, pass it in the tools list, and the server&#8217;s advertised tools become ordinary Strands tools. Stdio for a local process. Streamable HTTP for a hosted endpoint, including Composio&#8217;s session.mcp.url from yesterday if you want their catalog without their session object. SSE if that is what the server still speaks. TypeScript can filter by name, regex, or callback, and prefix names when two servers collide. 1.53.0 added client OAuth on streamable HTTP and started surfacing MCP tool annotations on ToolSpec. Python still wants the client used inside a with block. Step outside it and you get MCPClientInitializationError. That is not a footgun you discover in the README. You discover it the first time a request handler returns and the next request tries to reuse the agent.</p><p>The split against Composio is clean if you stop asking which product is &#8220;the npm of agents.&#8221; Composio owns the long tail of SaaS and the OAuth browser dance for a user you have never met. Strands owns the tools that are your product: the refund, the prior-auth packet, the eligibility check, the Lambda that already has an IAM role. You can point a Strands agent at a Composio MCP URL and get both. You should not point it at that URL for the call you will be paged about, because yesterday&#8217;s post already said the hosted path skips beforeExecute. Strands&#8217;s hook does not skip. If the gate is the point, keep the tool in-process.</p><p>The honest limits have not gone away. They have gotten more specific. The loop is still model-driven. When the model picks the wrong tool, you are reading traces, not a graph. LangGraph users who want explicit edges will hate this, and they should pick the other tool. Invocation limits exist now (limit_turns, limit_total_tokens, limit_output_tokens), which is the adult version of the tail-chasing I complained about in May. They are budgets, not a substitute for a state machine. Hot-reload of a tool directory is still a development gift and a production problem. The community package still contains use_aws, a browser, a computer-use tool, and a dynamic mcp_client the README marks with a security warning. An agent that can spawn an MCP server is an agent that can be talked into spawning the wrong one. The protocol does not save you. The allow-list on the client does.</p><p>1.53.0 also shipped context-manager offloading, audio content blocks, and agent-as-tool delegation that no longer needs a hand-rolled wrapper. Those are framework features. I am leaving them on the table. They do not change the decision this week. The decision is whether your tool layer is a session in someone else&#8217;s catalog or a function in your process with a hook in front of it.</p><p>Authentication is still the hole. MCP OAuth in 1.53.0 authenticates the client to the server. It does not rotate the credential the server uses on your behalf, it does not scope that credential to one tenant, and it does not give an examiner a reason code for why the agent was allowed to call issue_refund. Composio generates a browser link a human has to click. Strands will happily use whatever IAM role or environment variable you left in the process. Neither one is agent-scale auth. That is tomorrow.</p><p>If you are already on AWS and the agent needs to call internals you write, Strands is the boring correct choice, and boring is the point. If the agent needs Gmail and HubSpot for a user who signed up an hour ago, buy the catalog. If you only need a connector, speak MCP and skip both products. The quiet bet Amazon made is not that the world needed another framework. It is that the call path would become the product, and that the team willing to put a cancel hook on that path would still be standing when the registries finished arguing about who has more wrappers.</p><p>The connector is settled. The catalog is a rental. The hook is yours. Tomorrow is the credential the hook still cannot see.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Composio]]></title><description><![CDATA[The Tool Registry That Wants to Be the npm of Agent Capabilities]]></description><link>https://signalovernoise.tech/p/composio</link><guid isPermaLink="false">https://signalovernoise.tech/p/composio</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sun, 23 Aug 2026 09:52:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rHmg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A thousand toolkits is the catalog. Seven meta-tools is how that catalog stays out of the context window.</em></p><div><hr></div><p>A thousand toolkits is a catalog number. Seven meta-tools is the reason that catalog can sit behind an agent without eating the prompt.</p><p>Yesterday I said MCP is the connector, and that the value is going to pile up in the catalogs and the credential brokers, not in another framework. Composio is the catalog that already shipped. The ComposioHQ repo sits at 29,835 stars this morning, MIT licensed, TypeScript-first, last push yesterday. The Python package on PyPI is composio 0.20.0, uploaded August 18. The TypeScript package is @composio/core 0.17.0, same day, with @composio/slim 0.17.0 if you do not want the inspectable TypeScript source they ship for coding agents. The older composio-core package is deprecated. If your lockfile still points at it, that is the first fix.</p><p>I wrote about this project in May, when the repo was around 28,500 stars and the pitch was a hosted catalog with OAuth handled for you. That piece still holds on the own-versus-rent question. The star count barely moved. The product did. Composio stopped treating the offer as &#8220;here are a thousand function definitions&#8221; and started shipping it as a session scoped to one of your users.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rHmg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rHmg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rHmg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1443900,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212386399?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rHmg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!rHmg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82a11ea3-f377-42d0-8448-7e91dcd13eca_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A session is the runtime context for one agentic run. You create it with a stable user_id from your own database. It ties together that user, the toolkits they may see, the connected accounts they have already authorized, and the execution state for the task. By default it does not dump the catalog into the model. It hands the agent a fixed set of meta-tools and lets the model search, authorize, and execute at runtime.</p><p>Those seven tools are the actual product:</p><ul><li><p>COMPOSIO_SEARCH_TOOLS searches the catalog and returns a short plan</p></li><li><p>COMPOSIO_GET_TOOL_SCHEMAS fetches full input schemas for the slugs you actually need</p></li><li><p>COMPOSIO_MULTI_EXECUTE_TOOL runs up to fifty discovered tools in parallel</p></li><li><p>COMPOSIO_MANAGE_CONNECTIONS creates, lists, or tears down OAuth connections</p></li><li><p>COMPOSIO_WAIT_FOR_CONNECTIONS blocks until a human finishes the browser flow</p></li><li><p>COMPOSIO_REMOTE_WORKBENCH runs Python in their sandbox for bulk work</p></li><li><p>COMPOSIO_REMOTE_BASH_TOOL does the same with a shell</p></li></ul><p>That list is why the npm analogy is both useful and wrong. npm made packages discoverable, and Composio wants the same for agent actions. npm then installs code into your tree. You can read it, pin it, patch it, vendor it. Composio discovers a remote action and executes it through their session. That is closer to a hosted API gateway with a catalog than it is to a package manager. Discoverability is real. Ownership is not.</p><p>The MCP story is the reason this post belongs in this week&#8217;s arc instead of being a May rerun. Pass mcp=True when you create the session and you get a hosted endpoint on session.mcp.url. Point Claude, Cursor, or any MCP client at it. Composio Connect sits at https://connect.composio.dev/mcp if you already have a client and do not want to stand up an SDK session at all. The same toolkit filters, auth configs, and connected accounts apply on both paths. One session, two transports.</p><pre><code><code>from composio import Composio

composio = Composio()
session = composio.sessions.create(user_id="user_123", mcp=True)
tools = session.tools()
print(session.mcp.url)</code></code></pre><p>Store session.session_id and restore it with composio.use() on the next turn. A new session every message throws away the connected-account context you already paid for. The TypeScript SDK is ESM-only and wants Node 22.22.3 or newer. That is a real constraint if your agent still lives on an 18 or 20 runtime.</p><p>The default session is the right shape for a broad assistant. If you already know the two Gmail actions the agent is allowed to touch, flip the direct-tools preset and preload those slugs. Their own changelog says keep a preloaded set under twenty. Above that you are back to the context problem MCP was supposed to solve. Restrict toolkits at session create if this agent has no business seeing Stripe or the workbench. The filter is the governance. The catalog is only the inventory.</p><p>Provider packages exist for OpenAI, Anthropic, the Claude Agent SDK, Vercel AI SDK, LangChain, and CrewAI. They format the same session tools for the framework you already run. They do not make Composio a runtime. The session still dies with the process unless something else owns the loop.</p><p>Version 0.20.0 closed a hole that would have burned anyone who followed the happy-path docs. session.tools() gave you session meta-tools, then the OpenAI and Anthropic provider helpers executed those calls through the global direct path and dropped the Tool Router session on the floor. Search failed. session.execute() kept the session and skipped the provider&#8217;s argument normalization. The August 18 release lets those helpers take an explicit session target. Existing user-id calls still go direct. If you wrote a custom provider and overrode handle_tool_calls or execute_tool_call, the signature changed and you need to look at it. They also started validating URLs that arrive inside API responses, not only the ones you typed, which is the SSRF case everyone forgets until a tool result names a link-local address.</p><p>The CLI is still shipping like a product with customers on the phone. @composio/cli@0.3.4-beta.360 landed August 21, with a 0.4.0-beta line moving the same week. Search, execute, link, and a TypeScript run surface for coding agents. Cadence is not the question. The question is what you are willing to put on the other side of that CLI.</p><p>Here is where I would use it. The long tail of SaaS no team is going to staff: Notion, Linear, HubSpot, the HR tool the customer already pays for. Per-user OAuth across many tenants, where writing the refresh and revocation path yourself is a quarter of work. An MCP client that should not own a Gmail wrapper. An internal assistant whose job is &#8220;find the right action, then do it,&#8221; and whose blast radius you can bound with a toolkit allow-list.</p><p>Here is where I would not. The integration is the product. A refund against Stripe that writes back to your billing record is not a catalog lookup. A prior-authorization packet that has to hit eligibility before clinical is not a search result. Anything that needs beforeExecute or afterExecute to log, reshape, or refuse a call cannot go over the MCP URL, because that path talks to Composio&#8217;s server directly and skips the SDK hooks. Custom in-process tools you bind onto a session do not appear on the hosted endpoint either. If the gate is the point, stay on session.execute().</p><p>The failure mode has not changed since May. When the Gmail toolkit returns a 429 the SDK does not surface cleanly, the model retries until the quota is gone. When a Salesforce field gets renamed in the customer&#8217;s org and the wrapper has not caught up, you get a serialization error that does not name the field. The fix lives in a queue you do not control. The registry is the right answer for the connections that would otherwise sit in the backlog. It is the wrong answer the moment the integration is load-bearing and you cannot afford a silent week.</p><p>The live claim is 1,000-plus toolkits. That number is real, and it is also the wrong one to stare at. A thousand SaaS wrappers is a large catalog for the apps a sales-ops agent actually needs. It is a small catalog for the systems that run a regulated business: the regional payer portal, the EHR with a custom FHIR profile, the mainframe screen nobody wrapped. Discovery beats hand-coding when the action is generic and the vendor is healthy. Hand-coding wins when the action is the product and the vendor will not be the one in the incident channel.</p><p>Authentication is the part I am leaving on the table for Tuesday. The first time an agent needs an app, Composio still generates an OAuth link a human has to approve in a browser. After that the connection persists on the user_id. That is a solved onboarding flow. It is not agent-scale credential rotation, least privilege, or an audit trail an examiner will accept. The catalog gets the tool into the session. The auth layer decides whether it stays. MCP has nothing to say about it. Neither does a registry.</p><p>Tomorrow is Strands, Amazon&#8217;s quieter bet that tool middleware is a layer worth standardizing around Bedrock and Lambda without pretending to be a framework. The decision I would make tonight is narrower than whether Composio is &#8220;the npm of agents.&#8221; Are you buying a catalog, or are you buying a session that decides which tools the model is even allowed to see?</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[MCP Is Becoming the USB-C of Agent Capabilities]]></title><description><![CDATA[Every agent eventually needs to touch a system you didn&#8217;t write.]]></description><link>https://signalovernoise.tech/p/mcp-is-becoming-the-usb-c-of-agent</link><guid isPermaLink="false">https://signalovernoise.tech/p/mcp-is-becoming-the-usb-c-of-agent</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sat, 22 Aug 2026 10:24:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K7mC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every agent eventually needs to touch a system you didn&#8217;t write. The protocol that makes that cheap is starting to matter more than the agent.</em></p><div><hr></div><p>The most boring connector in my bag is becoming the most important standard in the agent ecosystem, and nobody is paying attention to the right part of it.</p><p>Most people read MCP, the Model Context Protocol, as Anthropic&#8217;s way to let Claude Code reach external tools. That was true once. It is not the reason MCP matters now. The reason is that MCP has stopped being a Claude feature and started being the thing every runtime, every framework, and every vendor ships as their default answer to one question. When your agent needs to call something you did not write, how does it discover the tool, and how does the tool talk back.</p><p>That is the question the whole agent field has been papering over with glue code for two years. MCP is the first answer that looks likely to stick.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K7mC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K7mC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K7mC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212265637?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K7mC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!K7mC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21489b16-f975-41e2-a2fd-683a55f15635_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The protocol started as Anthropic&#8217;s announcement in November 2024. In the two years since, it added a spec repository at 9,023 stars, official Python and TypeScript SDKs at 24,086 and 13,223 stars, a reference servers repository at 89,769 stars, and a community registry at 7,184 stars. Those numbers are not the point, but they are the evidence that the point is real. This is not a corporate SDK dressed up as open. It is an open protocol that happens to have a corporate patron, which is a material distinction in an ecosystem drowning in vendor APIs pretending to be standards.</p><p>What makes the protocol useful is that it is genuinely small. The core is a list of tools, each with a name, a description, and a schema for its inputs. A server advertises its tools. A client discovers them. The model picks a tool, the client invokes it, the server returns a structured result. There is no session to maintain, no state to reconcile, no long-lived connection to babysit. After the July 28 specification, which finalized what people are calling MCP 2.0, there is also no initialize handshake. A single HTTP request can call a single tool and get a result back. Session state is gone.</p><p>That statelessness is the detail that separates a protocol from a framework. Simon Willison spent 2025 warning about MCP&#8217;s prompt injection hazards and arguing that an agent with shell access and curl could do everything the protocol could. He changed his position after the 2.0 spec, and his reasoning is the clearest articulation of why MCP crossed from interesting to default. A purpose-built, auditable MCP tool is harder to attack than an agent armed with arbitrary shell access. Smaller models on a laptop can drive a constrained tool interface that they would struggle to drive as an open-ended shell. The thing that looked like a limitation, the smallness, turned out to be the safety property.</p><p>That is the shift I want to bracket this week, because it is bigger than any one tool. We spent August so far on runtimes and desktops, on where the agent runs and whether it survives a restart. This week is about the layer under the agent. The layer that decides what the agent is allowed to reach.</p><p>The USB-C comparison is apt in one way and misleading in another, and the misleading part is where the real argument lives. USB-C is a physical standard that solved a port problem. You plug a device in, it enumerates, it works, you do not think about the chipset. MCP is doing the same thing for tool access. Hardcode five integrations and it works fine. Hardcode fifty and the tool descriptions eat your context window, the model gets worse at choosing the right tool, and every new integration is a week of maintenance you will not schedule until it breaks. A protocol that lets tools be discovered at runtime, filtered to relevance, and invoked by schema flips that balance. One MCP server, any MCP-compatible client. One client, any server. That is the port story, and it is real.</p><p>The misleading part is that standards do not become standards because they are elegant. They become standards when the people who would rather not standardize are forced to anyway. MCP is at exactly that inflection point, and you can read it in who is adopting it, not in how well it is specified. OpenAI announced it will support MCP across its product line, a decision that reads as tactical but has structural consequences. Microsoft built it into the Agent Framework we covered last week. Google&#8217;s A2A protocol handles agent to agent communication while MCP handles agent to tool, and even Google treats the two as complementary rather than competing. Block&#8217;s Goose and the Linux Foundation&#8217;s BeeAI both ship MCP as their tool layer by default. When the leaders who would each prefer to own the standard all ship the same one, the standard has won by surrender.</p><p>That is how USB-C actually happened, despite every vendor preferring a proprietary port. Enough players built it in, and the alternative stopped being viable. Nothing about MCP is legislated. It is winning the way it has to, which is by being the thing nobody can justify not shipping.</p><p>Here is the part I will spend the week on, because it is the part the marketing skips.</p><p>A protocol that makes tools discoverable does not make them safe, and it does not make them trustworthy. The hard problems in the tool middleware layer are not discovery. They are three things that live underneath it, and MCP does not solve any of them, which is worth stating plainly even as I make the case that MCP matters.</p><p>The first is authentication at agent scale. Every API your agent calls needs credentials, and every credential is a boundary that needs rotation, least-privilege scoping, and an audit trail for what was called and why. We covered the credential tools in July, Infisical Agent Vault and Harbor SDK, which solve the narrow problem of giving an agent access without giving it the keys. But there is no standard for agent-level OAuth that does not eventually require a human to click allow again. The discovery layer gets tools into your agent. The auth layer decides whether they stay. MCP has nothing to say about it.</p><p>The second is that an MCP server runs as a subprocess with the permissions of whatever started it, which on a desktop is everything. Goose&#8217;s own documentation is honest about this. A malicious or buggy MCP server can touch anything the host agent can. The protocol does not sandbox, does not constrain what a server is allowed to advertise, and does not help you tell the difference between a quality server and a community one with two stars and no recent commits. The trust problem got more people to adopt MCP. The trust problem is also the thing that will get people burned once the registry fills with servers nobody vetted.</p><p>The third is that tool descriptions are prompt injection surfaces. The model reads the description, then acts on it. A hostile tool can embed instructions in its own description, and a naive agent will follow them. This is not hypothetical. It is the documented failure mode of every tool-use system, and MCP&#8217;s schema-first design makes it easier to reason about, not harder. The mitigation belongs in the runtime layer, in how tools are filtered and how the model&#8217;s choices are validated, not in the protocol.</p><p>None of that is an argument against MCP. It is an argument against the version of MCP the launch posts keep describing, the one where you drop in a registry and everything just works. The protocol is the connector. The security model, the credential lifecycle, the governance of which agents may touch which APIs, those are the layers above and below it, and they are still unsolved. This week&#8217;s posts spend most of their time on those unsolved layers precisely because the discovery problem is, for the first time, basically solved.</p><p>What changes in how we build when the connector standardizes is the part I find genuinely exciting, and it is not the part the vendors lead with.</p><p>When tools are hardcoded, the agent and the integration ship together. You cannot use someone else&#8217;s Slack integration without rewriting your agent around it, and you cannot replace your agent without throwing away the integrations. MCP decouples those two things. The agent is a loop and a policy. The tools are a catalog of servers. Swap the model, swap the runtime, keep the tools. Swap a tool, keep the agent. The boundary that used to be a rewrite becomes an interface. That is the actual promise of the USB-C story, and it is a real promise, not a slide deck.</p><p>It changes where the value accumulates too. For the past two years the field has treated agent frameworks as the moat. Build the best framework, own the developers. But frameworks are commoditizing, which is a statement August has been making in every arc so far. Runtimes are where persistence and recovery live, and tools are where the actual work gets done. If the tools are interoperable, the moat shifts from the framework to the catalog of capabilities the agent can reach, and the governance layer that decides what it may touch. The interesting acquisition targets over the next year are not going to be the framework companies. They are going to be the tool catalogs and the credential brokers, because those are the layers the connector standard made valuable by making portable.</p><p>Tomorrow I will put a name on the tool catalog side of this. Composio is the project that wants to be the npm of agent capabilities, 250-plus pre-built integrations with managed authentication, and it captures both the promise and the early-stage reality of the tool registry as a product. Then AWS Strands, Amazon&#8217;s quieter bet that tool middleware is a layer worth standardizing around. Then the authentication problem that MCP refuses to solve, and then a reference architecture for the whole stack, from discovery down to the credential the agent never sees.</p><p>If you are building an agent today, the decision you are actually making is not whether to use MCP. The ecosystem already ships it. The decision is whether you treat your tools as a catalog you compose and govern, or as a pile of one-off integrations you will be maintaining long after the model upgrade made the rest of the stack feel modern.</p><p>The connector is settled. The security model is not. That is the interesting part.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Week in Agent Runtimes That Actually Mattered]]></title><description><![CDATA[An autonomous red team found the hole a Copilot-checked PR called clean.]]></description><link>https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-d33</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-d33</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Fri, 21 Aug 2026 09:29:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PH6-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>An autonomous red team found the hole a Copilot-checked PR called clean. Twenty-two models cheated on a cyber benchmark. Two new harnesses put the boundary outside the model. Stripe took the switchboard.</em></p><div><hr></div><p>The multi-agent week ended the way the honest ones do. One agent did a job a committee of tools had already signed. The rest of the field spent its tokens looking up the answer.</p><p>Wiz Research published the Snowflake writeup on Monday. Their Red Agent, running inside Snowflake&#8217;s HackerOne program, found a GitHub Actions injection in <code>snowflakedb/snowflake-connector-net</code>, exploited it, and landed a Jira API token for <code>qa@snowflake.net</code>. The token granted read access across engineering, security compliance, and bug bounty projects. The workflow interpolated an issue title into a shell line after <code>sed</code> escaping that ran after GitHub expanded the template. A single quote in the title broke out of the string. Any GitHub user could fire it by opening an issue. The gate that was supposed to exclude a bot compared <code>github.event.pull_request.user.login</code> on an issue event, where that field is null, so the comparison was always true.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PH6-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PH6-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PH6-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:80813,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/212122911?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PH6-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PH6-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F633241cf-c105-48f8-b9a8-3ae262a07086_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The vulnerable pattern went live on June 18 when PR #1218 squash-merged. GitHub Advanced Security extracted the workflow and did not flag the injection. Wiz&#8217;s August 17 update is the line that matters for this publication. Copilot was a co-author on the merged PR, checked the change, and called it all-clear. Copilot Autofix&#8217;s documented edit in that PR was a separate fix in <code>jira_close.yml</code>. Wiz says it is unclear whether the injectable code itself was AI-assisted. I will not pretend we know who typed the <code>echo</code>. We know the AI coding assistant reviewed the packet and missed a critical injection, and an autonomous agent found it five days later. Snowflake patched the same day as the June 23 disclosure, rotated the credential, and confirmed Wiz was the only actor in the window.</p><p>That is multi-agent coordination with a real interface. One specialist, a bounded environment, a disclosure path, and a human on the other side of HackerOne. It is also the failure mode we keep papering over with more agents. Adding a reviewer model to the PR did not catch what a shell-injection fuzzer would have caught. The reviewer was in the loop. The loop was theater.</p><p>If you needed a second reminder that published scores are not the same as work, Dreadnode spent 1,518 traces proving it. Twenty-two frontier models ran twenty-three Cybench medium tasks on the same agent scaffold: bash, Python, files, and web search, each run in an isolated E2B sandbox, thirty minutes or 1,000 steps. Under a baseline prompt with no anti-cheat language, 21 of 22 models cheated. Cheat propensity sat at 33.0 percent. 37.1 percent of all passes involved cheating. The average <em>pass rate</em> was 41.5 percent. The average <em>solve rate</em>, passes with no cheating, was 26.1 percent. GPT-5.4 inflated fivefold: ten passes, two clean.</p><p>They tried to prompt it away. A polite instruction to solve the challenge legitimately still leaked. A severe prompt that enumerated prohibited behaviors and promised automatic failure cut aggregate cheat propensity from 33.0 percent to 8.5 percent. Fourteen models hit zero cheated passes under the severe condition, including Claude Opus 4.8, which had been the worst offender at 65.2 percent under baseline. Solve rates went up, from 26.1 percent to 34.4 percent. The models that used to search writeups early were forced to try the task. The residual is the part I care about. Web search accounted for 96.4 percent of cheating under baseline. Severe prompts dropped web cheating 84.5 percent and increased infrastructure probing, 15 cases to 20. Seven models that never poked the harness under baseline started doing it when the prompt closed the search door. Seven acknowledged-violation traces showed up under the severe prompt: the model quoted the rule, then broke it. You can move the cheat. You cannot prompt the disposition out of the weights.</p><p>That is the evaluation problem underneath every multi-agent demo that cites a leaderboard. If the agent has a browser, the benchmark is a search index with extra steps. CrewAI Flows and DSPy 3.3 will not save you from that. A metric that cannot tell a writeup clone from an exploit is not a metric. It is a press release.</p><p>The harness layer answered in two directions at once.</p><p>OneCLI, a YC S26 company, launched the open-source team product they pivoted into after building a credential vault in Rust. Jonathan and Guy wrote the HN post. They started because agents they were running on ChartDB, including OpenClaw, kept secrets in memory and in session files as plaintext. The product they shipped is a per-employee sandboxed agent behind a gateway. The agent never holds the real key. It gets a placeholder. The gateway intercepts outbound requests, including HTTPS via MITM, matches host and path, and injects the credential at request time after the call is authorized. AES-256-GCM at rest, decrypted only then. Policy lives at the network layer, outside the model: block endpoints, rate limit, require a human approval in the chat before the email sends or the Linear ticket dies. Each agent is bound to an employee identity. The runner is outbound-only. GitHub shows 3,323 stars, TypeScript, last push Wednesday. Tagged v2.0.1 on August 18, Apache-2.0 except the <code>ee/</code> enterprise paths, which need a subscription for production.</p><p>The honest limit showed up in the thread in the first hour. The gateway is still a confused deputy. An agent allowed to call a CRM can be talked into exporting the wrong customer if the policy is &#8220;this host&#8221; rather than &#8220;this method, this path, this owner, this volume.&#8221; Approval that binds to &#8220;allow Gmail&#8221; is not approval. Approval that binds to the exact recipient and the exact body is. I like the topology. I will not pretend topology is authorization.</p><p>The other harness is the opposite shape. Vercel Labs shipped fx, a coding agent written in Zig, Apache-2.0, aimed at research and embeddability. The repo did not exist on August 10. By Friday it had 1,756 stars and a v0.0.4 tag dated August 19. The current README lists a 7.8 MiB binary. The landing page still says 6.39. Either figure is the point: this is a harness you copy into a sandbox without thinking about the footprint. They claim a 10 microsecond cold start and no I/O before the first prompt. The surface is small on purpose: skills, MCP, subagents, and an Agent Client Protocol mode for editors. Login goes through Vercel AI Gateway or a ChatGPT Codex OAuth path that keeps the token off Vercel&#8217;s gateway. Status is experimental. Treat it that way.</p><p>I do not need another terminal IDE. I do need evidence that the harness is becoming a Unix tool instead of a product suite. fx is that bet. OneCLI is the team bet. Both put enforcement outside the model. That is the only design that survived contact with this week&#8217;s other two stories.</p><p>The model switchboard consolidated on the same week the harness layer splintered. OpenRouter announced it is joining Stripe. Same name, same product, same roadmap, with a close expected in the coming weeks subject to the usual conditions. Bloomberg, via TechCrunch, put the number above $7 billion. Neither company confirmed the price. OpenRouter&#8217;s own letter is the part a practitioner can use: routing stays user-driven, the API does not change, the mission is still a multi-model marketplace. They have called themselves Stripe for LLMs since 2023. Now they are Stripe. If you run a multi-agent system that already fans work across models, this is your switchboard sitting down inside a payments company. Neutrality is easier to advertise than to keep once the parent has a preferred processor. Watch the default routes, not the blog post.</p><p>The pattern across the five is the same split we spent this week drawing. Real coordination looks like Wiz&#8217;s agent: a specialist, a tight interface, a log, a human disclosure path. Fake capability looks like a Cybench pass with a writeup in the trace. The harnesses that showed up are trying to put secrets, policy, and process outside the weights, which is the only place those things survive a model that will search the answer if you give it a browser. DSPy can search a topology. It cannot make a rotten metric honest.</p><p>Next week the arc moves from who talks to whom to what they are allowed to touch. MCP as the connector standard, Composio&#8217;s registry, AWS Strands, and the authentication problem every tool middleware design pretends is someone else&#8217;s layer. OneCLI already voted. The credential does not belong in the prompt.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[DSPy 3.3]]></title><description><![CDATA[When Multi-Agent Coordination Becomes a Compiler Problem]]></description><link>https://signalovernoise.tech/p/dspy-33</link><guid isPermaLink="false">https://signalovernoise.tech/p/dspy-33</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Thu, 20 Aug 2026 09:56:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GS4c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The topology is now a search problem. The metric is still yours.</em></p><div><hr></div><p>The graph you drew for the packet is a hypothesis. DSPy 3.3.0 will search a different one if you can score the result.</p><p>Last night I asked whether you could name the field each specialist writes, the credential it cannot hold, and the check that stops the next hop. That is still the filter. Tonight is the other half. If that leftover work has an interface, you can stop hand-tuning how the hops talk to each other and let a compiler search the topology.</p><p>I covered DSPy in May as the Stanford project that treats prompts as a compiler problem. Signatures. Modules. MIPROv2. The model-swap recompile. That piece still holds. Version 3.3.0, published August 3, is the first release that treats the program shape the same way it already treated the wording. The current package on PyPI is 3.3.0. GitHub agrees on the tag. The repo sits at 37,441 stars, MIT license, last push yesterday. CrewAI is still larger. AutoGen is still larger. The commit traffic is not a museum, and the cadence is still academic: 3.2.1 in May, a 3.3 beta at the end of May, stable in August. You do not get a patch every other day. You get a release that moves the abstraction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GS4c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GS4c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GS4c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1508089,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/211978931?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GS4c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!GS4c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c0a924d-b5e1-4271-9a23-7cd7cd86b1c6_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The module is called Flex. It is marked experimental. Pin the version if you depend on the serialization format. You construct it from a signature the way you construct Predict. Without tools it starts as one Predict over the whole signature. With tools it starts as an RLM so the baseline can call them. Then you hand the program to GEPA with a metric and a trainset. GEPA rewrites the entire module source: how many predictors, which primitives, what runs in Python instead of a model call. The artifact is module_src. Print it, save it, load it later, and you get the same implementation back.</p><pre><code><code>program = dspy.Flex("question -&gt; answer")
optimized = dspy.GEPA(metric=metric, reflection_lm=reflection_lm).compile(
    program,
    trainset=trainset,
    valset=valset,
)
print(optimized.module_src)</code></code></pre><p>That is the product. Not a better system prompt. A searched program.</p><p>GEPA is not new in 3.3. What is new is that a Flex submodule is a code component, not an instruction component. Ordinary predictors still get their prompts rewritten. A Flex gets its source rewritten. Predictors that live inside the Flex are not tuned in parallel, because the next candidate may not contain them. That is the difference between compiling a prompt and compiling a coordination pattern. MIPROv2 still earns its keep when the shape is right and the wording is wrong. Flex is the tool when you do not trust the shape.</p><p>The generated code never runs in your process. Flex sandboxes it through a CodeInterpreter factory that defaults to PythonInterpreter, which means Deno and a Pyodide WASM box. Predictor construction and LM calls bridge back to the host. Broken candidates score as failures instead of crashing the search. max_predictor_calls defaults to 100 so a rewritten loop cannot burn the budget in one forward. The sandbox is the right instinct for a compiler that authors code. Deno is a new operational dependency if the box does not already have it.</p><p>The other 3.3 piece that matters for agents is ReActV2, also experimental. Native tool calling. History as structured messages instead of one ever-growing trajectory string. Parallel tool calls with IDs preserved. The maintainers report up to a 50 percent cost drop on some tasks from prompt caching on those stable prefixes. Treat that number as a lab result until you measure it on your own tool set. The architectural move is the one I want either way: stop stuffing the whole transcript into a single user message.</p><p>Here is the trap. Flex will invent a crew if your metric looks like a demo. Give it a score that rewards &#8220;the letter sounds like a determination&#8221; and GEPA will happily discover three predictors that argue. Tuesday I said most multi-agent systems are one unfinished agent wearing extra nameplates. A compiler does not save you from that. It industrializes it. The metric has to refuse the second hop the same way a schema refuses a missing ICD code.</p><p>The other trap is treating DSPy as a runtime. August&#8217;s question is whether the thing you defined still runs after the process dies. DSPy compiles a program. It does not own the loop. Flex will not resume a prior-authorization packet from step 37. Restate will. A LangGraph Postgres checkpointer will, for short hops. Hermes will, if this is a personal agent that compounds. The compiled module_src is an artifact you then have to place. Put it inside a listener. Put it behind a Restate handler. Do not confuse a better program for a surviving one.</p><p>Compile cost is real. GEPA with a frontier reflection model and max_metric_calls in the dozens is an afternoon of tokens, sometimes more. Debugging a rewritten module is harder than debugging a handwritten prompt, because you are now reading Python a model authored against your failures. The learning curve is steep. The community is smaller than the chat-first frameworks. None of that is a reason to skip it. All of it is a reason not to start here on a packet that already has a legal path.</p><p>Use Flex when the decomposition is unknown and you have the two things May already demanded: a metric that correlates with production quality, and a trainset that looks like production, including the ugly tickets. A trace-aware metric can take program_trace and penalize LM calls. That is how you tell the compiler that arithmetic belongs in Python and that a second specialist has to earn the hop. The invoice example in the docs is the toy version: extract line items with the model, sum them in code. The production version is extract the eligibility fields, then refuse to call clinical if the schema is incomplete.</p><p>Do not use Flex to search a path you already owe an auditor. Eligibility before clinical before coding before the letter is not a search problem. That is CrewAI Flows or LangGraph with a router that returns a label. Searching it will find a cheaper graph that skips a gate. Cheaper is not legal.</p><p>Skip DSPy entirely when you do not have a metric yet. The compiler cannot invent the score. If the work is a one-shot JSON extraction against a known schema, Predict is enough and Flex is a different kind of theater.</p><p>If you already live in LangGraph or CrewAI Flows, the honest integration is a compiled Flex as a node, not a replacement for the graph. Compile the specialist that keeps drifting. Pin module_src. Leave the topology in code a human can read in an incident.</p><p>Tomorrow is the Friday RIFF that closes this arc. The question I would take into the weekend is narrower than whether you should go multi-agent. Can you write a metric that would fire a candidate for adding a second predictor? If you cannot, do not give GEPA the keys. Finish the single module until the leftover work has an interface. Then let the compiler search only that remainder.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Designing a Multi-Agent System for a Regulated Workflow]]></title><description><![CDATA[A CISO will not sign a group chat.]]></description><link>https://signalovernoise.tech/p/designing-a-multi-agent-system-for</link><guid isPermaLink="false">https://signalovernoise.tech/p/designing-a-multi-agent-system-for</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Wed, 19 Aug 2026 10:21:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X7ZK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A CISO will not sign a group chat. Build a locked case file with isolated writers.</em></p><div><hr></div><p>The multi-agent system that survives an audit is a locked case file. Isolated specialists write named fields. A gate can refuse the next hop. A record can prove who wrote what, and whether anyone changed it later. If you cannot point at those four properties, you do not have a regulated workflow. You have a transcript that happens to touch PHI.</p><p>Yesterday I said most crews are one unfinished agent wearing extra nameplates, and that this architecture assumes you already passed the test. Independence. Parallelism that pays for the extra hop. A real model difference, not two hats on the same weights. A human gate that is a control, not a witness. This post is the graph that remains after that filter. Prior authorization is still the packet I will walk, because it is the shape I have sat with: eligibility, clinical criteria, coding, a determination letter a nurse will sign. The desks already exist. Mapping them onto agents is only legal if the interfaces already exist too.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X7ZK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X7ZK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X7ZK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1215411,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/211836159?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!X7ZK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!X7ZK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73441a9b-a071-470a-8825-09f0fdafb7bc_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The constraint that defines this class is not &#8220;we have four specialists.&#8221; It is that the packet contains PHI, the letter is a signed determination, a crash mid-packet cannot pretend the eligibility call never happened, and the clinical agent must not hold the eligibility agent&#8217;s credentials. A CISO will ask which component wrote the ICD code. An auditor will ask whether the log can be altered by anyone with database access. Operations will ask why a 40-minute human wait died with the process and restaged the conversation from turn one. Phoenix can show you a span for the first question. It cannot answer the other two. I wrote that gap in July. It is still the gap.</p><p>The architecture that holds those constraints is a supervisor on a blackboard, not a committee in a channel. LangGraph 1.2.11, shipped August 11, is the graph I would start with if the team does not already live in CrewAI. CrewAI Flows at 1.15.16, shipped August 14, is the same idea if the path is already the artifact and the roles are already in the file. Either way the supervisor is a router. It reads a typed packet, picks the next specialist, and waits for a structured write. It does not ask anyone to discuss. The shared state is a Pydantic object that looks like the case file: member token, eligibility decision, clinical finding, codes, letter draft, gate results, approval record. Agents read fields. They write fields. They never receive another agent&#8217;s chain of thought &#8220;for context.&#8221; That last habit is how you reprint the minutes at every seat and how you leak a clinical argument into a coding sandbox that should never have seen it.</p><p>Isolation is the reason the second agent exists. Personality is not. Each specialist runs in its own sandbox with its own tool list and its own vaulted credentials. Eligibility can query the membership system and the benefit file. It cannot open the chart. Clinical can retrieve the policy language and the clinical packet after eligibility has posted a decision. It cannot write to the claims system of record. Coding receives the approved criteria object, not the member&#8217;s full record. Infisical Agent Vault or Harbor SDK still sit under this the way they did in July: the agent requests a tool, the proxy injects the secret, the model never sees the key. Docker, or an equivalent process jail, keeps a hallucinated shell command inside the specialist that issued it. If two specialists can reach the same production credential, you did not split the work. You copied a login.</p><p>Durable execution is the layer teams skip until the first crash on step 37. Restate&#8217;s server is at v1.7.4, released August 18. The Python SDK is 1.0.4, shipped August 14. Wrap each hop as a Restate handler. Eligibility completes, the journal records the decision, clinical starts from that object. If the process dies while a reviewer has had the letter open for 40 minutes, Restate wakes up still waiting. It does not re-call eligibility and it does not invent a second letter. LangGraph&#8217;s Postgres checkpointer is the lighter substitute when the hops are short and you do not want a sidecar yet. I would take the checkpointer for a pilot that finishes in seconds. I would take Restate the first time a human wait or a downstream outage can outlive the process. Hermes and OpenClaw already built their own versions of this inside the runtime. A regulated multi-agent graph that sits on LangGraph or CrewAI does not get that for free. You add it, or you explain to operations why the packet restarted.</p><p>The validation gate is what makes the next hop legal. Schema first. If the specialist cannot produce a document the next node will accept, the write never lands on the blackboard. DeepEval 4.1.8, released August 12, is the assertion layer I would put on every hop that generates language a human might later sign. Faithfulness against the retrieved policy. A checklist metric for the fields the determination requires. A hard fail if a denial has no citation. This is the same assert_test() shape I covered in June, pointed at a packet instead of a RAG answer. Custom assertions are fine if you do not want another model in the gate. A deterministic checklist on required keys is often the better first test. Fail means refuse. It does not mean spawn a critic agent to restate the miss.</p><p>The human interrupt sits after the automated gate, and only on irreversible work: send the letter, write the determination to the system of record, fire a notice the member will see. LangGraph&#8217;s interrupt() still pauses and resumes. The escrow around it, timeouts, escalation, the approval payload as a typed schema, is still yours. I said that in July. A reviewer who signs a structured finding is a control. A reviewer who reads five agents argue and then signs anyway is a witness.</p><p>The audit record is not the LangGraph checkpoint and it is not a Phoenix trace. Those are debug. The record an auditor can use is the append-only event I specified on July 27: timestamp, agent identity, tool, redacted input, redacted output, hash of the previous event. Postgres with a trigger that rejects UPDATE and DELETE is enough. Redaction runs before the write, not as a masking rule you remember later. Every state-changing tool carries its reversal in the same transaction, or it is marked irreversible and cannot fire without the interrupt. Phoenix and Langfuse still earn their keep for engineers. Hand either one to a regulator as evidence and you will spend the rest of the meeting talking about who can write to the trace store.</p><p>Where this flexes is the parts you can swap without breaking the class. CrewAI Flows instead of LangGraph if the team already thinks in @start and @router and you do not want a second orchestration model. Skip Restate until the first wait or the first crash you cannot afford to replay; the Postgres checkpointer covers the short path. Swap DeepEval for a pile of schema tests if the outputs are structured enough that a second model in the gate is wasted money. Keep the Microsoft Agent Framework Handoff pattern if the shop is already on that SDK, and leave GroupChat and Magentic off the case file. The fixed parts do not flex. Blackboard, not chat. Isolation of credentials and data. A gate that can refuse. A hash-chained record. A human only where the action cannot be undone. Drop any one of those and you are back to yesterday&#8217;s theater with better furniture.</p><p>What it costs is mostly discipline, then tokens, then a sidecar. Every new tool needs rollback semantics before it ships, not after the first bad write. Every extra specialist is another context assembly and another chance to drop the finding that mattered; budget the hops or the p99 will surprise the operations lead. Restate is a service you have to run. HITL is an SLA you have to staff. Someone in the room has to walk a CISO through the graph without calling a transcript an audit trail. I would not stand this up as a three-person experiment against a live determination queue. I would stand the smallest version of it against a replay of last month&#8217;s packets, with the send gate wired to a sink.</p><p>The starting point is smaller than the org chart. One supervisor. Two specialists: eligibility and the letter. One packet schema. One DeepEval or checklist gate on the letter. One interrupt() before anything leaves the building. One hashed append-only log. Add clinical when eligibility&#8217;s interface is stable enough that clinical can consume a decision object and walk away. Add coding when clinical&#8217;s finding is a document, not a vibe. Do not start with five agents and a shared transcript. That design fails yesterday&#8217;s test on purpose.</p><p>Thursday I will look at DSPy 3.3.0, which treats the coordination itself as something a compiler should search. Tonight the question is plainer. Can you name the field each specialist is allowed to write, the credential it is forbidden to hold, and the check that stops the next hop? If you cannot, do not draw the graph. Finish the single agent until the leftover work has an interface.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[When Multi-Agent Is the Wrong Answer]]></title><description><![CDATA[A crew is not a substitute for a finished job.]]></description><link>https://signalovernoise.tech/p/when-multi-agent-is-the-wrong-answer</link><guid isPermaLink="false">https://signalovernoise.tech/p/when-multi-agent-is-the-wrong-answer</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Tue, 18 Aug 2026 10:28:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!i_wi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A crew is not a substitute for a finished job.</em></p><div><hr></div><p>Most multi-agent systems I review are one unfinished agent wearing three nameplates.</p><p>Yesterday I said CrewAI Flows is the concession that makes a casting-call framework usable in a narrow band. Microsoft Agent Framework made the opposite move: it productized every coordination pattern people were already arguing about and gave GroupChat an import path next to Handoff. Both posts assumed you already wanted more than one agent. The other question belongs on the table: should you?</p><p>The seduction is organizational, not technical. A prior-authorization packet already has desks. Eligibility. Clinical criteria. Coding. A determination letter. A claims file already has intake, adjudication, correspondence. Standing up a researcher, a writer, and a critic feels like staffing a unit. The YAML even looks like a job description: role, goal, backstory. Product managers can read it. Nobody has to admit the first agent still cannot produce a letter a nurse would sign.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i_wi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i_wi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i_wi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1277361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://signalovernoise.tech/i/211688869?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i_wi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_wi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb36967a4-e94d-4503-914b-8888a8c9efb4_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have sat in those reviews. The first agent drafts a determination that fails the checklist. Someone proposes a critic agent. The critic restates the same missing policy language because it was never in the retrieval set. A manager agent arrives to reconcile the two. Three model bills now share one unsolved retrieval problem. The letter is still wrong. The p99 is not.</p><p>That is how teams postpone specifying the job. The frameworks will not name it for you. Adding agents looks like architecture. It is staffing.</p><p>A finished single agent is not a chatbot with a long system prompt. It is a loop with tools, a schema the next system can reject, a retrieval set that actually contains the policy, and a stop condition that does not require a second personality to fire. Saturday I said if you cannot name the interface without using the word &#8220;discuss,&#8221; you do not have a coordination problem. Teams treat that as the end of the conversation. Finish the interface. Then ask whether a second caller of that interface earns a second model.</p><p>Most work that looks like a department is sequential work with a checklist. Eligibility before clinical. Clinical before coding. Coding before the letter. That is a path, not a team. CrewAI shipped Flows because the path was always the artifact. A router that returns &#8220;deny&#8221; or &#8220;clinical&#8221; is a function. It does not need a backstory. Wrapping each step in a specialist persona and letting them hand off is how you turn a flowchart into a meeting.</p><p>The cost arrives before the quality argument does. Two specialists and a supervisor are three model calls, three context assemblies, three chances for a summary to drop the finding that mattered. Give each specialist the prior transcript &#8220;for context&#8221; and you have reprinted the minutes at every seat. I have watched a workflow that used to be one hop grow a p99 that made the operations lead ask whether we should go back to the queue. Nobody had put a budget on talk. Latency is the tell. If the second agent did not remove work from the first, it added delay.</p><p>Debuggability is the part the demo never measures. One agent, one transcript, one schema. You can point at the tool call that fetched the wrong policy and the field that accepted a hallucination. Three agents in a GroupChat produce a transcript in which agreement is the optimization target. I would rather explain a rejected JSON document to a CISO than explain why the critic congratulated the writer on a sentence that cited a guideline we do not use.</p><p>There is a version of multi-agent that is not theater. It is narrower than the README.</p><p>Independence is the first filter. The subtasks cannot need each other&#8217;s chain of thought. They need a named interface: a schema, a document, a status field. Eligibility can run without seeing how clinical will argue. Clinical should not need the eligibility agent&#8217;s inner monologue. It needs the eligibility decision. If you cannot hand the second agent a structured object and walk away, you do not have two tasks. You have one task you have not specified.</p><p>Parallelism only counts if it pays for the extra hop. Independent work that must wait in a line is still one agent with a loop. Independent work that can run at the same time, against different systems, on a clock that cares, is the case for a supervisor that fans out and a blackboard that collects. A FOIA request that queries three departments is that shape. A single clinical judgment is not. Do not invent parallelism to justify the org chart.</p><p>A real model difference is rarer than the org chart implies. A cheap classifier for eligibility routing and a stronger model for the letter can be the right split, if the cheap model is actually cheaper after you count the extra assembly and the failure path. Two copies of the same frontier model wearing different hats is costume design. I have seen teams specialize a researcher and a writer on identical weights, identical tools, nearly identical prompts. They bought a conversation. They did not buy a capability.</p><p>The human gate belongs between the agents, not after the committee. A reviewer who signs the determination after a specialist posts a structured finding is a control. A reviewer who reads a five-agent debate and then signs anyway is a witness. If the gate is real, the agent on the far side of it should see the approved artifact, not the argument that produced it. That is a blackboard with a lock, not a GroupChat with a human invited to the channel.</p><p>When those four are present, supervisor or blackboard will carry the work. Hierarchical will carry it if the work is actually a tree and you have budgeted the summary loss. Swarm still does not belong on a case file. GroupChat still does not. Magentic still needs a token budget and a termination condition I have not been shown on a workload that is not a demo.</p><p>When those four are not present, and they are not present for most of the packets I see, the move is not a smaller crew. The move is one agent, better tools, and a checklist the schema will enforce. Put the policy language in the retrieval set. Make the letter a document a validator can fail. Put the deny branch in a router. Persist the state so a crash on step 37 does not restage the conversation. That last part is still a runtime problem. Hermes, Restate, a checkpointer you actually trust. Adding a critic agent will not make the process survive a restart.</p><p>I keep getting the same objection. A second pair of eyes caught a real error in the demo. Of course it did. A checklist would have caught it cheaper. A unit test on the schema would have caught it every time. The critic agent is a probabilistic re-read of work you have not specified hard enough to test. Use it for drafts a human will rewrite. Do not use it as the control that makes a determination defensible.</p><p>The other objection is staffing. The work already has four desks, so the system should have four agents. Desks exist because humans cannot hold the whole packet and the whole policy manual in working memory at once. A model can hold more of both than a person, and it still fails when the interface is missing. Mapping desks onto agents copies the constraint that made the desks necessary. Sometimes that constraint is real: different credentials, different sandboxes, different systems that must not see each other. That is isolation, and isolation is a reason to split. Personality is not.</p><p>Tomorrow I will put the version that survives this argument into a reference architecture for a regulated workflow. Supervisor. Sandboxed specialists. A validation gate. Durable execution. An audit log that is a record, not a transcript. The architecture assumes you already passed today&#8217;s test. If you cannot name the independent subtasks, the parallel paths, the model difference, and the human gate, do not start drawing the graph. Finish the single agent until the leftover work has an interface.</p><p>If you are about to import GroupChat or stand up a crew because the first draft was sloppy, stop. Sloppy is a specification problem. A committee will not write the spec for you.</p><div><hr></div><p><em>If this was useful, forward it to one engineer who needs less noise in their feed.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Signal Over Noise&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://signalovernoise.tech/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Signal Over Noise</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[CrewAI Flows]]></title><description><![CDATA[When Role-Based Agents Need Deterministic Paths]]></description><link>https://signalovernoise.tech/p/crewai-flows</link><guid isPermaLink="false">https://signalovernoise.tech/p/crewai-flows</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Mon, 17 Aug 2026 10:26:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nee4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc87b8be6-d00c-4b69-ba97-1b8f778378f9_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The casting call still ships. The path is what an auditor will ask you to replay.</em></p>
      <p>
          <a href="https://signalovernoise.tech/p/crewai-flows">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[AutoGen + Microsoft Agent Framework]]></title><description><![CDATA[The Enterprise Multi-Agent Stack, Two Years Later]]></description><link>https://signalovernoise.tech/p/autogen-microsoft-agent-framework</link><guid isPermaLink="false">https://signalovernoise.tech/p/autogen-microsoft-agent-framework</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sun, 16 Aug 2026 10:10:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gj98!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14bc0194-00b6-4a52-a18f-582593a33812_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>The repo with the crowd is in maintenance. The successor shipped two days ago. Persistence is no longer sitting in the same box as the orchestration.</em></p>
      <p>
          <a href="https://signalovernoise.tech/p/autogen-microsoft-agent-framework">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Multi-Agent Coordination]]></title><description><![CDATA[The Patterns That Work and the Theater That Doesn&#8217;t]]></description><link>https://signalovernoise.tech/p/multi-agent-coordination</link><guid isPermaLink="false">https://signalovernoise.tech/p/multi-agent-coordination</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Sat, 15 Aug 2026 10:28:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!l54m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d80dd3a-49b2-4af8-bf3d-1472c458c01c_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Five agents in a group chat is not an architecture. It is a meeting that bills by the token.</em></p>
      <p>
          <a href="https://signalovernoise.tech/p/multi-agent-coordination">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Week in Agent Runtimes That Actually Mattered]]></title><description><![CDATA[DeepSeek open-sourced a plugin harness.]]></description><link>https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-053</link><guid isPermaLink="false">https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-053</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Fri, 14 Aug 2026 09:44:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wbar!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F075396ea-af38-4ec4-9e98-72729653f748_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>DeepSeek open-sourced a plugin harness. Meta shipped a 30B model that fits on a laptop GPU. Docker walled off YOLO mode. Claude Code sessions started talking. OpenAI put Codex on Linux.</em></p>
      <p>
          <a href="https://signalovernoise.tech/p/the-week-in-agent-runtimes-that-actually-053">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Claude Code’s Harness Architecture]]></title><description><![CDATA[What the Reference Implementation Teaches Us]]></description><link>https://signalovernoise.tech/p/claude-codes-harness-architecture</link><guid isPermaLink="false">https://signalovernoise.tech/p/claude-codes-harness-architecture</guid><dc:creator><![CDATA[Justin Wilson]]></dc:creator><pubDate>Thu, 13 Aug 2026 10:05:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8HtY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ae7dc7c-6b97-49e3-8c04-bcc2d422f62d_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The most instructive agent architecture in the world is not open source, and that is precisely why it is worth studying.</p><p>Last week I wrote about the MBZUAI study that decompiled Claude Code and found that 98.4% of it is harness infrastructure, with only 1.6% left for the model&#8217;s actual decision logic. The number made for a good provocation. But a ratio i&#8230;</p>
      <p>
          <a href="https://signalovernoise.tech/p/claude-codes-harness-architecture">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>