Blog
AI Agents
Building, deploying, and managing autonomous AI agents for workflow automation, customer support, internal operations, and more.
642 articles in this category
Apollo Asked If an Agentic Bank Run Is Coming. The Question Is Who Runs the Agent.
Apollo chief economist Torsten Sløk asked on 27 September 2026 whether agentic AI assistants could sweep household cash out of 0.1% checking accounts into the 3.3% to 5.0% accounts his note lists. The mechanism he describes needs an agent holding account access, and the bank that does not operate that agent does not get a vote in what it optimizes for.
Bad Theory Labs' Interference Search: Check the Scope
Bad Theory Labs published a paper, Apache-2.0 code, trained judge weights, raw results and a log mapping every number to the command that produced it. The figures are real measurements — on Countdown, an arithmetic puzzle with an exact solver. The unsupported step is not the lab's; it is the leap from that benchmark to enterprise agent reasoning.
Microsoft's Agent Identity Layer: Who Actually Owns It?
Microsoft's rebuilt Copilot puts its persistent Autopilot agent behind a governed identity and meters it with usage-based billing rather than the per-seat licence. The governance is good and the metering is the right shape. What you cannot get is ownership of the control plane, which is the part worth pricing before you standardize on it.
You Cannot Govern a Clinical Model You Cannot Observe
Anthropic reported that ~950 agents surfaced a novel enzyme system in 21 hours, and the first FDA-approved AI margin-assessment device reached its first operating rooms in August. The capability question is closing. The governance one is not — and even that approved device ships AI updates under a change-control plan the hospital does not hold.
Forward-Deployed Engineering: The Four-Layer Agent Stack
In every enterprise agent deployment we have worked on, the same four layers appear: an environment provisioner, an evaluation harness, a governed deployment stack, and simulation. This page publishes them as an ownership checklist — which layer you hold, which your vendor holds, and what breaks when the answer is your vendor.
Nubank Screened 16,000 Simulated Chats Before Going Live
Nubank screened open-weight model configurations across more than 16,000 simulated conversations before putting one in front of customers, and published the results. The paper's buried finding is not the simulation count — it is that a bank serving 140 million customers in a regulated market chose an open-weight model, and simulation is what made that choice defensible.
AI Agents Do Licensed Work. Liability Law Doesn't Fit.
Agents now draft motions and reason across clinical documents — work that is licensed when a human does it. Product liability assumes a defect, professional liability assumes a licensed practitioner, and agency law reaches non-human agents only partway. At least seven states have answered by prohibiting AI therapy; the Cicero Institute proposes licensing the service instead.
When Your AI Defender and Your AI Threat Share a Vendor
In one UN General Assembly week OpenAI gave Ukraine its Daybreak cyber-defence system free, and Australia revealed that an OpenAI agent had circumvented access controls on a Medicare statistics portal in June. A day later the disclosures widened again. Throughout, the detection function sat with the vendor.
Agent Containment Moved Into Silicon. What You Still Own.
NVIDIA's Open Agent Safety Platform pairs OpenShell, an Apache 2.0 runtime boundary you can run today, with Sentry — an out-of-band watchdog NVIDIA describes as a reference system design, not a shipping product. The strongest control in the announcement is the one you cannot buy yet, and that is the part worth planning around.
Model-Agnostic Is Table Stakes. Source Code Isn't.
Microsoft Foundry has documented model-agnostic agents for months, and Microsoft even ships a disconnected on-premises agent path. Both are real, and neither is the thing enterprises should be buying on. The question a managed service still cannot answer is whether you receive the source code.
Two Decision Models in Nine Days. One You Can Own.
TypeSafe shipped Jev on 15 September 2026 and Fastino shipped GLiNER2.5-Decide on 24 September. Two vendors, nine days, the same architectural claim: the routing and classification an agent does all day should not run on a frontier model. One of the two is Apache 2.0.
Letting a K-12 AI Agent Run Code Without Letting Data Out
An AI tutor that can actually run a student's Python is worth far more than one that can only talk about it — but only if the district can say where that code ran and what it could reach. Agent sandboxes answer that with a Linux VM that starts with no network at all.
Why PII Redaction Breaks Enterprise AI — and What Fixes It
Redaction removes the relationships that made enterprise data worth training on. Transformation models keep them by swapping real identities for consistent synthetic ones — and runtime filtering catches the PII that arrives after training, in chat, in uploads, in screenshots.
Who Audits the AI Writing Into the Nurse's Flowsheet?
Oracle Health made its Clinical AI Agent for nurses available in the US on September 14, 2026, for structured documentation at the bedside, in fields that sit outside FDA device oversight.
An Agent That Moves Money Needs Row-Level Permissions
AWS open-sourced TOLAP on 22 September 2026 under Apache-2.0: row filtering and column masking enforced inside agent tools, across three SDKs and fourteen framework integrations.
Whose Agents Run Your Operations? The Managed Agent War
Five vendors now sell managed enterprise agent platforms, from AWS Bedrock AgentCore in October 2025 to Google's Gemini Enterprise Agent Platform in April 2026. All of them operate the agents for you, and that is the whole buying decision.
Agent Infrastructure Went Open Source. The Record Didn't.
Chutes and researchers at Harvard and Chicago released 6,122,413,756 production LLM requests across 9,174 models — in twelve metadata fields that hold no prompts, no responses and no tool calls.
Distributing Your Agent Platform Changes What You Own
Pine Labs and Google Cloud announced Gemini-powered merchant agents on 24 September 2026, serving over 1 million merchants on ₹17.15 trillion of FY26 transaction value, and distribution is what changes the ownership question.
Hospitals Buy the Charting Bot, Not the Early Warning
Documentation AI sells because its benefit is visible on day one; sepsis early warning saves more lives but the saved death is a counterfactual. Medicare now pays up to $61.84 per eligible case for it.
YC Open-Sourced QM: The Harness Became Infrastructure
Y Combinator open-sourced QM on 31 July 2026 under an MIT license, and says it runs the harness internally across four departments — not across its portfolio. The substance is scoping.
Agent Memory Fails at State Transitions, Not Retrieval
On StateMemBench, released August 2026, leading agent memory systems answered with the current state only 13–20% of the time. The failure is not retrieval quality — nothing marks a fact as superseded.
Beijing's AI Eye Clinic Hit 3.8% Clinician Adoption
A Nature Medicine Comment published 10 September 2026 reports that Beijing Tsinghua Changgung Hospital's AI-TEC agent clinic was used in 41 of 1,113 examinations — 3.8% — before workflow changes lifted it to 23%.
When Your Software Vendor Applies for a Bank Charter
Block applied on 8 September 2026 for an uninsured national trust bank that cannot take deposits or lend, and the OCC has taken 40 de novo charter applications in 18 months against 48 in the 14 years to 2024.
Generation Is Commoditized. Judgment Is the New Frontier
TypeSafe announced Jev on September 15, 2026 — a decision model priced at $0.042 per million input tokens with output unmetered. It is not the first model built to judge: CriticGPT and Prometheus 2 both shipped in 2024.
About AI Agents
AI agents represent the next evolution in enterprise automation—intelligent systems that can reason, plan, and take action autonomously. Unlike simple chatbots, AI agents handle complex multi-step tasks across customer support, internal operations, data analysis, and specialized workflows. Discover how agentic AI is transforming how organizations operate.