AI Devtools Daily — Friday, July 31, 2026
Yesterday the US only petitioned for agent governance. Today the binding version already exists abroad, the enterprise plumbing to let agents act on systems of record went live, and an open-weight model matched the frontier on SWE-bench.
Three threads this briefing has tracked all month resolved into facts today. On governance: yesterday's US "Pacing the Frontier" letter was a voluntary ask, but the enforceable version is already law elsewhere — China's three-tier agent-authorization framework (human-only, approval-required, autonomous) became binding on July 15, and Illinois now mandates third-party AI safety audits, creating dual compliance deadlines for any enterprise agent team. On interoperability: MCP became the enterprise lingua franca — Google, Microsoft, Salesforce, Snowflake, and ServiceNow all now support it, ServiceNow opened its full "system of action" to any agent, and Snowflake is acquiring MCP-governance startup Natoma. On the frontier: open weights caught up — Meituan's LongCat-2.0 (1.6T MoE, 48B active, 1M context) edges GPT-5.5 on SWE-bench Pro and claims parity with Gemini 3.1 Pro on software engineering, fully open. And the money confirmed where value is pooling: Databricks raised ~$3B at a $188B valuation, explicitly for agent-workload data infrastructure (Lakebase, Unity AI Gateway). The durable businesses are the layers that make agent authority governable, MCP access controllable, open-weight parity provable, and agent writes to systems of record reversible.
TL;DR
- Binding agent law now exists — abroad. China's three-tier decision-authorization framework became enforceable July 15; agents in sensitive sectors face filing, mandatory testing, and recall. Illinois added third-party safety audits. The US petitioned; China legislated.
- MCP is the enterprise standard. Google, Microsoft, Salesforce, Snowflake, and ServiceNow all back it; ServiceNow opened its system of action to any agent; Snowflake is buying Natoma for MCP governance. The protocol war is over — the governance war is starting.
- Open weights hit the frontier on code. LongCat-2.0 (1.6T open MoE) edges GPT-5.5 on SWE-bench Pro and claims Gemini-3.1-Pro-class SWE performance. Sovereign, self-hosted, frontier-class coding is now feasible.
- Agents got write access to systems of record. ServiceNow opening its full system of action — plus China's human-override mandate — makes staged, reversible agent writes a requirement, not a nicety.
- Money pooled in agent-workload data infra. Databricks' ~$3B round at $188B targets Lakebase and Unity AI Gateway — the transactional and gateway layers agents actually run on.
Market trends
Agent governance went from a voluntary US ask to binding foreign law.
Yesterday's "Pacing the Frontier" letter asked Washington to help. But the enforceable version already shipped: China's Implementation Opinions (effective July 15) sort every agent action into three tiers of decision authority and require documented autonomy limits and human-override mechanisms before deployment, with filing, testing, and recall for sensitive sectors. Illinois separately mandates third-party AI safety audits. For any enterprise running agents across markets, governance is no longer a policy debate — it is a dated compliance obligation with more than one jurisdiction to satisfy.
Pebblous — China's three tiers · MachineBrief — enforceable July 15 · AI Governance Institute — dual deadlines
MCP won the protocol war — now the governance war begins.
Google, Microsoft, Salesforce, Snowflake, and ServiceNow all supporting MCP makes it the enterprise interoperability standard. ServiceNow opened its full system of action to any MCP agent; Snowflake is acquiring Natoma specifically to govern agent access. Once every vendor exposes an MCP server, the hard question is not "can agents connect" but "which agent may call which server, with what scope, logged how" — and each vendor governs only its own surface. The cross-vendor access-control plane is the open gap.
CIO — Snowflake acquires Natoma · ServiceNow — system of action to every agent · Microsoft Azure — MCP + A2A open standards
Open weights reached the frontier on agentic coding.
LongCat-2.0 — 1.6T total parameters, ~48B active, native 1M context, open — reportedly edges GPT-5.5 on SWE-bench Pro and claims Gemini-3.1-Pro-class software-engineering performance. Whatever survives independent testing, the direction is clear: a self-hostable model is now within striking distance of the proprietary frontier on code. For regulated, air-gapped, and sovereignty-constrained buyers who cannot send their codebase to a US cloud, frontier-class coding just became an on-premises option.
MarkTechPost — LongCat-2.0 1.6T open MoE · DeepWiki — agentic & coding benchmarks · NxCode — routing, benchmarks, cost
Agents now write to systems of record — so writes need an undo.
ServiceNow opening its full system of action to every agent means agents move from reading and drafting to committing changes in the enterprise's authoritative systems. China's law simultaneously requires human override above set thresholds. The blast radius of a bad agent write just grew from a wrong answer to a corrupted record of truth. Staged approval and rollback for agent actions on systems of record shifts from optional guardrail to operational and legal requirement.
ServiceNow — full system of action · Pebblous — human-override thresholds · AI Agents News — week of July 30
Fresh product / business ideas
Writ
a decision-authority tiering engine for enterprise agents
A policy layer that classifies every agent action into human-only, approval-required, or autonomous — enforces the boundary at call time, routes approval-required actions to a human, and keeps a tamper-evident log auditors can read — because China's binding three-tier framework and Illinois' audit mandate just made documented autonomy limits a legal requirement, not a design choice.
- Who it’s for
- Enterprises running agents in China, Illinois, or any jurisdiction adopting tiered-authority rules, plus multinationals that must satisfy several at once.
- Why now
- China's framework became enforceable July 15 and Illinois added third-party audits — the obligation is dated and live. Distinct from Cordon (July 28, technical sandbox-escape detection) and Hallmark (July 30, deployed-model capability attestation): Writ governs business decision authority per action, not model internals.
- First version
- An SDK/proxy that intercepts agent tool calls, maps each to a tier via configurable policy, blocks or escalates the approval-required set to a human queue, and emits a signed, per-action authority log with jurisdiction tags.
- What kills it
- Harnesses ship native approval gates. Counter: those are per-tool and unauditable across a fleet — the cross-tool, jurisdiction-aware, evidence-producing authority engine is the compliance artifact a single harness will not build.
Concourse
a vendor-neutral MCP access-governance plane
One control plane for which agents may call which MCP servers across Google, Microsoft, Salesforce, Snowflake, and ServiceNow — identity, scope, rate, and a unified audit trail — because MCP just became the enterprise standard and Snowflake bought Natoma to govern its own surface, leaving no one governing the space between vendors.
- Who it’s for
- Platform and security teams whose agents now span five-plus vendor MCP ecosystems, each with its own siloed access model.
- Why now
- Five major vendors backing MCP and Snowflake acquiring Natoma (this week) make cross-vendor MCP sprawl a present reality. Distinct from Bridgehead (July 28, stateless-spec compatibility shim): Concourse is authorization and audit across servers, not protocol translation.
- First version
- A gateway that brokers every agent-to-MCP-server call, applies per-agent scopes and rate limits, and produces a single audit stream across all connected vendor servers.
- What kills it
- A vendor extends its governance across rivals. Counter: no vendor will neutrally govern access to a competitor's server — neutrality is the whole product.
Bastion
a sovereign, self-hosted agentic-coding appliance on open weights
A turnkey on-premises coding-agent stack built on LongCat-2.0-class open weights, shipped with an eval harness that proves parity with Opus/Gemini on the customer's own task set — because open weights just reached the frontier on SWE-bench and regulated, air-gapped, and sovereignty-bound shops still cannot send code to a US cloud.
- Who it’s for
- Defense, government, and IP-sensitive organizations that need frontier-class coding but must keep every token inside their perimeter.
- Why now
- LongCat-2.0 edging GPT-5.5 on SWE-bench Pro (open weights) makes self-hosted parity credible for the first time. Distinct from Homestead (July 14, managed local-inference fleet) and Ballast (July 27, self-host-vs-API break-even calculator): Bastion is the deployable sovereign coding appliance plus a defensible parity proof, not advice or a general fleet.
- First version
- A hardened deployment (model + agent harness + repo-context server) for air-gapped clusters, bundled with a benchmark that scores the open model against frontier APIs on the buyer's tasks and outputs a defensible parity report.
- What kills it
- Hosted vendors ship compliant sovereign tiers. Counter: "in our cloud" is exactly what these buyers cannot accept — fully in-perimeter with independent parity evidence is the wedge.
Throughline
longitudinal agent-behavior regression testing
A monitoring service that tests whether a deployed agent stays consistent across sessions — does it respect standing user preferences, drift in behavior, or quietly change how it handles the same request over weeks — because VitaBench 2.0 just made long-term, multi-session agent behavior the thing worth measuring, and no one runs that check in production.
- Who it’s for
- Teams running long-lived, memory-equipped agents where week-over-week consistency matters more than one-shot accuracy.
- Why now
- VitaBench 2.0's shift to long-term personalized/proactive evaluation (this week) reframes quality as longitudinal, not per-response. Distinct from Marginalia (July 30, the memory store itself): Throughline tests the behavior that memory produces over time.
- First version
- A harness that replays a golden set of multi-session scenarios against a live agent on a schedule, flags drift and preference violations, and trends consistency over time.
- What kills it
- Eval vendors add longitudinal suites. Counter: most eval tooling is one-shot and offline — continuous, in-production, cross-session drift detection is a different product.
Backstop
staged-write approval and rollback for agent actions on systems of record
A layer that intercepts agent writes to ServiceNow, Salesforce, and Snowflake, stages them for approval above a configurable risk threshold, and makes every committed change reversible with a one-click undo and full lineage — because ServiceNow just opened its system of action to every agent and China's law now requires human override above set thresholds.
- Who it’s for
- Enterprises letting agents act on authoritative systems, who now face a blast radius that reaches the record of truth.
- Why now
- ServiceNow opening its full system of action plus binding human-override rules (this week) turn reversible agent writes into a requirement. Distinct from Writ (#1, which decides authority): Backstop handles the write itself — staging, approval, and rollback of the committed change.
- First version
- A write-interceptor for the major systems of record that queues high-risk changes for human sign-off, records before/after state, and offers atomic rollback with an audit trail.
- What kills it
- Systems of record add native agent-write approval. Counter: each covers only its own system — the cross-system staged-write and rollback plane spanning ServiceNow, Salesforce, and Snowflake is the layer none of them owns.
Worth watching
- How many jurisdictions adopt tiered agent-authority rules — if China's three-tier model spreads and Illinois-style audits multiply, decision-authority governance (Writ #1) becomes a category, not a one-country compliance chore.
- The MCP governance land-grab after Snowflake–Natoma — whether other vendors buy or build MCP access control, and whether a neutral cross-vendor plane (Concourse #2) gets room before incumbents close it.
- Independent verification of LongCat-2.0's SWE claims — whether open-weight parity holds outside Meituan's benchmarks, and how fast sovereign buyers move (Bastion #3).
- ServiceNow system-of-action adoption — how quickly agents get production write access to systems of record, and what breaks first (Backstop #5).
- Databricks' agent-workload data bets — whether Lakebase and Unity AI Gateway define the transactional and gateway layers agents standardize on.
Sources
- Pebblous — China three-tier agent rules
- MachineBrief — enforceable July 15
- AI Governance Institute — dual deadlines
- CIO — Snowflake acquires Natoma
- ServiceNow — system of action to every agent
- Microsoft Azure — MCP + A2A
- MarkTechPost — LongCat-2.0
- DeepWiki — LongCat-2.0 coding benchmarks
- AI Agents News — week of July 30
- llm-stats — AI news