Commentary

Auto Added by WPeMatico

Who Authorized That? The Delegation Problem in Multi-Agent AI

Your AI agent booked a meeting, summarized a financial report, and emailed the highlights to three stakeholders. To do this, it called a calendar agent, a document analysis agent, and an email agent. Each accessed internal systems, made decisions about what to include, and acted on your behalf. Here’s the question your security team can’t

Who Authorized That? The Delegation Problem in Multi-Agent AI Read More »

Figure 1. Empire of headcount vs. federated nervous system—An analogy

The Agentic P&L: Beyond the Empire of Headcount

For over a century, both the prestige and budget of a corporate department have been measured by a single crude metric: headcount. If you manage 500 people, you’re a “distinguished leader.” If you manage five, you’re a footnote. This “empire of headcount” has governed everything from office square footage to C-suite influence. It’s the fundamental

The Agentic P&L: Beyond the Empire of Headcount Read More »

Figure 1. Empire of headcount vs. federated nervous system—An analogy

The Agentic P&L: Beyond the Empire of Headcount

For over a century, both the prestige and budget of a corporate department have been measured by a single crude metric: headcount. If you manage 500 people, you’re a “distinguished leader.” If you manage five, you’re a footnote. This “empire of headcount” has governed everything from office square footage to C-suite influence. It’s the fundamental

The Agentic P&L: Beyond the Empire of Headcount Read More »

When an Agent Deletes the Production Database

Another day, another example of an AI Agent “running rogue” and doing something the human operator didn’t want it to do. The tl;dr is that Jeremy (Jer) Crane, founder of PocketOS, was using Claude to perform some routine DB maintenance. Claude then proceeded to delete the production database and all backups hosted at their cloud

When an Agent Deletes the Production Database Read More »

A minimalist catalog stored as a Git repository

AI Artifact Catalogs: Durable Standards Worth Institutional Investment

Companies everywhere are trying to leverage AI to boost internal productivity metrics. Some, like Ramp and Intercom, are succeeding. Many are failing. To make matters more complicated, the narrative around what tooling enables these gains is constantly shifting. For software engineers, auto-complete via GitHub Copilot was the bleeding-edge tool of choice in 2024. Then it

AI Artifact Catalogs: Durable Standards Worth Institutional Investment Read More »

The model is one chip on the board. The harness is everything else that makes it useful.

This article was originally published on Addy Osmani’s blog. It’s being reposted here with the author’s permission. Roughly: Anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again. We’ve spent the last two years arguing about models. Which one is

Read More »