Agents at work

Internal Agents Map

AI systems organizations build or adapt to do work for their own teams.

Explore how teams connect models to their knowledge, tools, and workflows—and where people stay involved. The map includes agents, platforms, orchestration systems, and supporting patterns.

39approaches
35organizations
88sources

Latest entry review: . Individual source dates vary.

Explore the catalog

Approaches, not rankings

39 approaches

AirbnbPlatform

Airchat (airchat-cli)

Airbnb's internal agentic-coding harness, built by its Dev AI team. Airchat is a wrapper over Claude Code with a unified gateway for cost and metrics, an internal plugin marketplace, AirDev Workspaces for parallel sessions, and more than a dozen internal MCP servers that connect agents to internal systems. The team abandoned an earlier from-scratch orchestrator and shipped a thin shim over Airchat instead.

CodingCode review
coding task → reviewed pull request: Work product review
Operating model, claims & sources

Scoped operating models

coding task → reviewed pull requestWork product review · Level 3

Summary and context

Summary

Airbnb's internal agentic-coding harness, built by its Dev AI team. Airchat is a wrapper over Claude Code with a unified gateway for cost and metrics, an internal plugin marketplace, AirDev Workspaces for parallel sessions, and more than a dozen internal MCP servers that connect agents to internal systems. The team abandoned an earlier from-scratch orchestrator and shipped a thin shim over Airchat instead.

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
The build is described by Airbnb engineers in talks and podcasts, not in a first-party engineering blog.

Reported metrics

Headline claim

About 64% of pull requests materialized through agentic coding

Metric · Reported · Low confidence
Evidence and qualifications
Confidence reason
The 64% figure comes from a third-party newsletter that quotes the engineers, not from a first-party Airbnb source.
Reported by
Airbnb
Scope
Airbnb PRs materialized through agentic coding by the October 2025 talk
Denominator
Airbnb pull requests; exact count not supplied
Method
Unknown
Observation date
2025-10
Key observation

About 64% of pull requests materialized through agentic coding

Metric · Reported · Low confidence
Evidence and qualifications
Confidence reason
The figure comes from a third-party newsletter, not a first-party Airbnb source.
Reported by
Airbnb
Scope
Airbnb PRs materialized through agentic coding by the October 2025 talk
Denominator
Airbnb pull requests; exact count not supplied
Method
Unknown
Observation date
2025-10

Architecture and primitives

Harness

Wrapper over Claude Code with a unified gateway, an internal plugin marketplace, and AirDev parallel workspaces

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Model

Claude Code (a vendor agent), wrapped by Airbnb

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Tool access

More than a dozen internal MCP servers connect agents to internal systems

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.

Operating model evidence

Operating model assessment

Level 3 for coding task → reviewed pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
Airbnb engineers describe agents producing pull requests that engineers review, which locates human attention at work-product review.
Observation date
2025

Sources

  1. Agentic coding at Airbnb (DPE.org) · Preserved Markdownhttps://dpe.org/sessions/szczepan-faber-mike-nakhimovich/agentic-coding-at-airbnb/talk · direct-participant · Last source verification: 2026-08-31
  2. Beyond the CLI (DX podcast) · Preserved Markdownhttps://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/podcast · direct-participant · Last source verification: 2026-08-31
  3. How to get your team past the AI (The AI Thinker) · Preserved Markdownhttps://www.theaithinker.com/p/how-to-get-your-team-past-the-ainews · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
AtlassianTask agent

Rovo Dev (RovoDev)

Atlassian's internal coding agent, built on the HULA (Human-in-the-loop software development agents) framework. Rovo Dev works inside Jira and runs a four-step cycle (set context, generate a plan, generate code, and raise a pull request). Atlassian dogfooded it across all Jira sites for more than a year across 1,900+ repositories. It reached general availability in October 2025.

CodingCode review
Jira issue → reviewed pull request: Work product review
Operating model, claims & sources

Scoped operating models

Jira issue → reviewed pull requestWork product review · Level 3

Summary and context

Summary

Atlassian's internal coding agent, built on the HULA (Human-in-the-loop software development agents) framework. Rovo Dev works inside Jira and runs a four-step cycle (set context, generate a plan, generate code, and raise a pull request). Atlassian dogfooded it across all Jira sites for more than a year across 1,900+ repositories. It reached general availability in October 2025.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Dogfooded across 1,900+ repositories with a 50,000+ comment internal dataset

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Atlassian reported the dogfooding scale in its own engineering blog without independent verification.
Reported by
Atlassian
Scope
Internal code-review dogfooding repositories and Rovo Dev-generated classifier training comments
Denominator
Unknown
Method
Unknown
Observation date
2026
Key observation

Dogfooded across 1,900+ repositories over more than a year

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Atlassian reported the dogfooding scale in its own engineering blog.
Reported by
Atlassian
Scope
Internal code-review evaluation across 1,900+ repositories spanning over a year
Denominator
Unknown
Method
Year-long online evaluation across internal repositories
Observation date
Unknown
Key observation

Trained on a proprietary internal dogfooding dataset of 50,000+ Rovo Dev comments

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Atlassian reported the dataset size in its own engineering blog.
Reported by
Atlassian
Scope
ModernBERT classifier training dataset of internally sourced Rovo Dev-generated comments
Denominator
Unknown
Method
Internal dogfooding comments labeled by whether they led to a code resolution
Observation date
Unknown

Architecture and primitives

Harness

Built on the HULA framework, which runs set context, generate plan, generate code, and raise PR

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Operating model evidence

Operating model assessment

Level 3 for Jira issue → reviewed pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The HULA framework and engineering blog describe a human-in-the-loop cycle that ends in a reviewed pull request.
Observation date
2024-11

Sources

  1. Improving the coding agent experience · Preserved Markdownhttps://www.atlassian.com/blog/atlassian-engineering/improving-coding-agent-experienceengineering-blog · first-party · Last source verification: 2026-08-31
  2. Developer productivity improved with Rovo Dev · Preserved Markdownhttps://www.atlassian.com/blog/atlassian-engineering/developer-productivity-improved-with-rovo-devengineering-blog · first-party · Last source verification: 2026-08-31
  3. HULA: Human-in-the-loop software development agents (arXiv 2411.12924) · Preserved Markdownhttps://arxiv.org/abs/2411.12924paper · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
BlockOrchestration system

Builderbot

A multi-agent orchestration layer built on goose + MCP that coordinates agents across Block's entire codebase, invoked in Slack to take a ticket end-to-end to a reviewed PR.

CodingCode review
ticket → reviewed pull request: Work product review
Operating model, claims & sources

Scoped operating models

ticket → reviewed pull requestWork product review · Level 3

Summary and context

Summary

A multi-agent orchestration layer built on goose + MCP that coordinates agents across Block's entire codebase, invoked in Slack to take a ticket end-to-end to a reviewed PR.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

~1,500 PRs merged per week (~15% of all production code changes at Block)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Block
Scope
Builderbot merged pull requests per week at Block
Denominator
All production code changes across Block for the approximately 15% share
Method
Unknown
Observation date
Unknown
Key observation

~1,500 pull requests merged per week (~15% of all production code changes at Block)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Block
Scope
Builderbot merged pull requests per week at Block
Denominator
All production code changes across Block for the approximately 15% share
Method
Unknown
Observation date
Unknown

Architecture and primitives

Knowledge

Company-wide code context across hundreds of millions of lines and hundreds of services; Block also frames Builderbot as an 'agentic protector' around its software world model

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Other reported details and interpretation

Lessons and interpretation

Operating model evidence

Sources

  1. Block rolls out Builderbot, a new suite of AI-native tools that changes the way we ship · Preserved Markdownhttps://block.xyz/inside/block-rolls-out-builderbot-a-new-suite-of-ai-native-tools-that-changes-the-way-we-shipengineering-blog · first-party · Last source verification: 2026-08-31
  2. Protecting our systems with intelligence (Builderbot architecture) · Preserved Markdownhttps://engineering.block.xyz/blog/protecting-our-systems-with-intelligenceengineering-blog · first-party · Last source verification: 2026-08-31
  3. block/builderbot source repository · Preserved Markdownhttps://github.com/block/builderbotrepository · first-party · Last source verification: 2026-08-31
  4. Hacker News discussion of the Builderbot announcement · Preserved Markdownhttps://news.ycombinator.com/item?id=48618973hn-thread · community · Last source verification: 2026-08-31
  5. Hacker News comment that points to the Builderbot repository · Preserved Markdownhttps://news.ycombinator.com/item?id=48619052hn-comment · community · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
BrexPlatform

Internal Agent Platform

Retool-based internal platform where employees build, test, and deploy agents for KYC, disputes, QA, collections, and operations.

Finance opsSupportCustomer success
internal operations request → completed operation: Continuous steering
Operating model, claims & sources

Scoped operating models

internal operations request → completed operationContinuous steering · Level 2

Summary and context

Summary

Retool-based internal platform where employees build, test, and deploy agents for KYC, disputes, QA, collections, and operations.

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.

Reported metrics

Headline claim

Dispute processing time fell from three hours to three seconds

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Brex
Scope
Dispute-submission preparation using the internal agent platform; not end-to-end chargeback resolution
Denominator
Unknown
Method
Unknown
Observation date
2025
Key observation

50%+ of customer-support cases resolved by chatbot as first touch

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Brex
Scope
Customer-support cases resolved by the chatbot at first touch
Denominator
Customer-support cases; exact sample size not provided
Method
Unknown
Observation date
2025
Key observation

Dispute processing: 3 hours → 3 seconds

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Brex
Scope
Dispute-submission preparation using the internal agent platform; not end-to-end chargeback resolution
Denominator
Unknown
Method
Unknown
Observation date
2025
Key observation

QA covers every support interaction; one person using AI instead of five QA specialists

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Brex
Scope
Quality assurance of customer-support interactions
Denominator
Every support interaction
Method
Agent applies the quality rubric to every response; one person oversees instead of five QA specialists
Observation date
2025
Key observation

KYC adverse-media accuracy 85% → 88%

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Brex
Scope
KYC adverse-media classification; human versus agent accuracy
Denominator
Unknown
Method
Unknown
Observation date
2025

Architecture and primitives

Harness

Retool-based builder with prompt management and multi-model testing/evaluation; built by a ~25-person systems-engineering team

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Tool access

An MCP server exposes external product features to the internal platform; new product tools become internally available immediately; invoked via Slack /c1

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Credentials

SSO via internal Retool proxies (no per-user accounts); ConductorOne access management; Okta auth; data classified by risk; ≤30-day retention, no training on inputs

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.

Lessons and interpretation

Operating model evidence

Operating model assessment

Level 2 for internal operations request → completed operation; human attention boundary: continuous-steering.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The secondary report describes a human remaining in the middle of the workflow, but does not fully specify each review surface.
Observation date
2025-09-25

Sources

  1. Agent, Human, Ops: How Brex Is Changing Roles and Workflows · Preserved Markdownhttps://www.firstround.com/ai/brexcase-study · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
BrowserbaseTask agent

bb

One generalized agent in Slack that writes PRs, investigates sessions, queries the warehouse, logs feature requests, and runs browser agents across engineering, ops, sales, and support.

CodingCode reviewSupportCustomer successResearch
coding request → reviewed pull request: Work product review
Operating model, claims & sources

Scoped operating models

coding request → reviewed pull requestWork product review · Level 3

Summary and context

Summary

One generalized agent in Slack that writes PRs, investigates sessions, queries the warehouse, logs feature requests, and runs browser agents across engineering, ops, sales, and support.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Feature-request pipeline at 100% coverage with zero human effort

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Browserbase
Scope
Automatic feature-request scanning of every closed support ticket and meeting transcript
Denominator
Closed support tickets and meeting transcripts
Method
Unknown
Observation date
Unknown
Key observation

Feature-request pipeline at 100% coverage, zero human effort

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Browserbase
Scope
Automatic feature-request scanning of every closed support ticket and meeting transcript
Denominator
Closed support tickets and meeting transcripts
Method
Unknown
Observation date
Unknown
Key observation

99% of first-response times < 24 hrs

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
The source reports 99% below 24 hours with awkward wording; no sample, timestamps, or method is provided.
Reported by
Browserbase
Scope
Reported support first-response time below 24 hours
Denominator
Support first responses; exact sample and measurement window not provided
Method
Unknown
Observation date
Unknown
Key observation

Session investigation: 30–60 min of log-diving → one Slack message

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A Slack message replaces the manual initiation workflow; the article does not report end-to-end automated investigation latency.
Reported by
Browserbase
Scope
Session investigation initiation: manual log-diving versus a Slack request
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Tool access

exec routes through a serverless integration proxy (Snowflake, HubSpot, Pylon, Grafana); the sandbox never sees real secrets

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Credential brokering; the sandbox boots with references + rotating session tokens only; the proxy holds real creds; egress injection for a few hosts

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

One agent with good abstractions beats a fleet of narrow bots

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Separate capabilities from the core loop; domain logic lives in skills + service packages

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Don't trust the model, remove its ability to do wrong: scope tools and services per invocation source

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Meet people where they are; Slack is the highest-leverage surface because that's where work already happens

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for coding request → reviewed pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source documents agent-authored pull requests while the enclosing workflow retains human review.
Observation date
2026

Sources

  1. How we build internal agents at Browserbase · Preserved Markdownhttps://browserbase.com/blog/internal-agentsengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
CloudflarePlatform

Internal AI engineering stack

An internal platform of MCP servers, an access layer, and AI tooling (incl. an AI code reviewer) that makes agents useful inside Cloudflare.

CodingCode review
pull request → AI review findings: Work product review
Operating model, claims & sources

Scoped operating models

pull request → AI review findingsWork product review · Level 3

Summary and context

Reported metrics

Headline claim

47.95 million AI requests in 30 days across the internal AI engineering system

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Cloudflare
Scope
Internal AI engineering requests in the 30 days preceding the report
Denominator
Unknown
Method
Company-reported AI Gateway count for the preceding 30 days
Observation date
2026
Key observation

3,683 internal users (60% of company, 93% of R&D)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Cloudflare
Scope
Active internal AI coding-tool users in the preceding 30 days
Denominator
Approximately 6,100 employees for company share; R&D organization for R&D share
Method
Unknown
Observation date
2026
Key observation

47.95M AI requests and 241.37B tokens via AI Gateway in the preceding 30 days

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Cloudflare
Scope
Internal AI requests and AI Gateway tokens in the preceding 30 days
Denominator
Unknown
Method
Reported request counts and AI Gateway token counts
Observation date
2026
Key observation

10,952 merge requests in the week of March 23, 2026, nearly double the Q4 baseline; four-week average above 8,700

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Cloudflare
Scope
Company merge requests in the week of March 23, 2026, versus Q4 baseline; not agent-authored PRs
Denominator
Unknown
Method
Weekly merge-request count; distinct from the four-week rolling average
Observation date
2026-03
Key observation

295 teams using agentic AI tools

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Cloudflare
Scope
Teams using agentic AI tools and coding assistants in the reported 30-day snapshot
Denominator
Unknown
Method
Unknown
Observation date
2026

Architecture and primitives

Tool access

MCP Server Portal; one OAuth point aggregating 182+ tools from 13 servers; AI Gateway for routing, cost, BYOK, ZDR

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Zero API keys on client machines; a Worker injects keys server-side; Cloudflare Access (Zero Trust) auth

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Code Mode collapses upstream tool schemas into search + execute, holding token overhead constant at scale

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Structured, generated repo context (runtime, nav, conventions, boundaries, deps)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Centralize through a proxy early; direct-to-gateway looks simpler but blocks per-user attribution, model cataloging, and policy later

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Without structured data, agents are working blind; they read code but can't see the system around it (Backstage / AGENTS.md)

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Tool schemas eat context (34 GitLab tools ≈ 7.5% of a 200K window); collapse them at the portal

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Frontier + open-source hybrid: route a growing share of workloads to cheaper self-hosted models

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for pull request → AI review findings; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source documents automated review findings, but the record combines several platform workflows.
Observation date
2026-04-20

Sources

  1. The AI engineering stack we built internally · Preserved Markdownhttps://blog.cloudflare.com/internal-ai-engineering-stack/engineering-blog · first-party · Last source verification: 2026-08-31
  2. Hacker News discussion of Cloudflare's internal AI engineering stack · Preserved Markdownhttps://news.ycombinator.com/item?id=47837240hn-thread · community · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
CoinbaseAgent system

Forge / Mux

Forge turns a Slack/GitHub/Linear discussion into a Linear issue, fix, PR, and one-off build; Mux lets employees run many coding agents concurrently.

CodingCode review
Slack, GitHub, or Linear request → reviewed pull request and build: Work product review
Operating model, claims & sources

Scoped operating models

Slack, GitHub, or Linear request → reviewed pull request and buildWork product review · Level 3

Summary and context

Reported metrics

Headline claim

Mux: 600+ users including engineers, PMs, and designers (335 active, 197 power users)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Coinbase
Scope
Registered Mux users including engineers, PMs, and designers; 335 active and 197 power users
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Mux: 600+ users including engineers, PMs, and designers (335 active, 197 power users)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Coinbase
Scope
Registered Mux users including engineers, PMs, and designers; 335 active and 197 power users
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Tool access

Linear treated as the structured product context / source of truth

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Linear as the durable structured-context layer for product work

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Custom harness: Slack bug discussion → Linear issue → fix → PR → one-off build

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Keep the system of record (Linear) as the agent's structured context; conversation can be the input, Linear stays durable

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Simple per-agent worktree/branch/terminal isolation is a pragmatic alternative to full sandboxing for concurrent coding

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for Slack, GitHub, or Linear request → reviewed pull request and build; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes delegated implementation that returns a pull request and build for human review.
Observation date
2026

Sources

  1. Coding had a concurrency problem: how Mux helped solve it · Preserved Markdownhttps://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-itengineering-blog · first-party · Last source verification: 2026-08-31
  2. Coinbase × Linear (Forge workflow) · Preserved Markdownhttps://linear.app/customers/coinbasecase-study · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
DatabricksAgent system

coSTAR and internal engineering agents

Databricks' internal engineering agents and the coSTAR framework that ships and tests them. Databricks uses internal agents as daily coding drivers on its own codebase, including code-review and on-call support work. coSTAR tests agents on a private benchmark built from Databricks' multi-million line codebase before they ship. Omnigent is a separate shipping open-source product and is excluded from this record.

CodingCode reviewOn-call
internal engineering workflows → agent-produced changes: Unknown
Operating model, claims & sources

Scoped operating models

internal engineering workflows → agent-produced changesUnknown · Level unknown

Summary and context

Summary

Databricks' internal engineering agents and the coSTAR framework that ships and tests them. Databricks uses internal agents as daily coding drivers on its own codebase, including code-review and on-call support work. coSTAR tests agents on a private benchmark built from Databricks' multi-million line codebase before they ship. Omnigent is a separate shipping open-source product and is excluded from this record.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Other reported details and interpretation

Operating model evidence

Operating model assessment

Unclassified for internal engineering workflows → agent-produced changes; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The record covers several internal engineering agents with different workflows, so no single human-attention boundary applies.
Observation date
2025

Sources

  1. coSTAR: how we ship AI agents at Databricks fast · Preserved Markdownhttps://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-thingsengineering-blog · first-party · Last source verification: 2026-08-31
  2. Benchmarking coding agents on a multi-million line codebase · Preserved Markdownhttps://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebaseengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-08-31
DomuTask agent

Clementino

A general-purpose 'AI colleague' spanning sales, finance, client ops, engineering, and recruitment, later split into a reusable toolkit plus Slack and desktop surfaces.

SupportFinance opsCodingRecruitmentCustomer success
employee request → approved customer-impacting action: Work product review
Operating model, claims & sources

Scoped operating models

employee request → approved customer-impacting actionWork product review · Level 3

Summary and context

Summary

A general-purpose 'AI colleague' spanning sales, finance, client ops, engineering, and recruitment, later split into a reusable toolkit plus Slack and desktop surfaces.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

~35 integrations organized into skills

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Domu
Scope
Clementino toolkit integration modules organized into skills
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

~35 integrations organized into skills

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Domu
Scope
Clementino toolkit integration modules organized into skills
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Harness

Claude/Anthropic SDK wrapper; a reusable tool/skill/memory layer separated from the Slack and desktop interfaces

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

~35 integrations organized into skills; specialist delegates per domain

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Four memory layers: conversation context, persistent facts, knowledge RAG, and live system state

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Customer-impacting actions gated behind team-visible Slack approvals

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Four-layer memory separation; prompt size dropped substantially after splitting capabilities from the Slack layer

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Transient conversation, persistent facts, knowledge RAG, and live system state; kept separate

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Capabilities separated from the interface layer so they power multiple surfaces

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Separate reusable capabilities (tools/skills/memory) from the interface layer; prompt size drops while capability is preserved

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Model memory explicitly: conversation context, persistent facts, RAG, and live state have different lifecycles

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Gate customer-impacting actions behind human approvals rather than trusting the model

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for employee request → approved customer-impacting action; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source documents explicit human approval for customer-impacting actions.
Observation date
2026

Sources

  1. Why we split our internal agent in two · Preserved Markdownhttps://domu.ai/blog/why-we-split-our-internal-agent-in-twoengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
DoorDashBackground agent

AI Code Review Agent

A specialized agent that automatically reviews 10,000+ PRs a week across 56 repositories, emphasizing grounded high-confidence findings over noisy comments.

Code review
pull request → AI review comments: Work product review
Operating model, claims & sources

Scoped operating models

pull request → AI review commentsWork product review · Level 3

Summary and context

Reported metrics

Headline claim

10,000+ pull requests reviewed per week across 56 repositories

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
DoorDash
Scope
Typical weekly PR reviews across 56 onboarded repositories
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

10,000+ PRs reviewed in a typical week across 56 repositories

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
DoorDash
Scope
Typical weekly PR reviews across 56 onboarded repositories
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

60.2% action rate on settled high/critical findings (measured sample)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
DoorDash
Scope
Settled high and critical findings that led to code changes before merge
Denominator
2,256 settled high and critical findings
Method
Whether the human changed the code before merge in response to the finding
Observation date
Unknown

Architecture and primitives

Lessons and interpretation

Operating model evidence

Operating model assessment

Level 3 for pull request → AI review comments; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes automated findings that engineers evaluate within the pull-request workflow.
Observation date
2026

Sources

  1. How DoorDash built an AI code reviewer engineers actually listen to · Preserved Markdownhttps://careersatdoordash.com/blog/doordash-built-an-ai-code-reviewer-engineers-actually-listen-to/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
DoorDashPlatform

Flux / Agentic AI Platform

DoorDash's internal agentic AI platform; a unified cognitive layer over company data and operations, with an AI Marketplace of specialized agents and the Flux cloud-agent runtime for engineering tasks.

Code reviewCodingCI triageOn-callMaintenanceData
engineering task → reviewed agent output: Work product review
Operating model, claims & sources

Scoped operating models

engineering task → reviewed agent outputWork product review · Level 3

Summary and context

Summary

DoorDash's internal agentic AI platform; a unified cognitive layer over company data and operations, with an AI Marketplace of specialized agents and the Flux cloud-agent runtime for engineering tasks.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

130,000 engineering tasks automated in one month

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Dated August 11, 2026 report of a one-month count; this is the observation date, not the measurement window.
Reported by
DoorDash
Scope
Engineering tasks automated in one reported month; calendar measurement month unspecified
Denominator
Unknown
Method
Unknown
Observation date
2026-08-11
Key observation

130,000 engineering tasks automated in one month

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Dated August 11, 2026 report of a one-month count; this is the observation date, not the measurement window.
Reported by
DoorDash
Scope
Engineering tasks automated in one reported month; calendar measurement month unspecified
Denominator
Unknown
Method
Unknown
Observation date
2026-08-11
Key observation

25,000+ automated code reviews per week

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Dated August 11, 2026 report; weekly measurement boundaries are not supplied.
Reported by
DoorDash
Scope
Weekly automated code reviews powered by Flux
Denominator
Unknown
Method
Unknown
Observation date
2026-08-11
Key observation

300+ playbooks; 10,000+ invocations per week

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
DoorDash
Scope
Unique playbooks and weekly invocations on Flux
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Lessons and interpretation

Lesson

Start narrow to earn trust; began with automated code review before CI triage, on-call, maintenance, ticket-driven dev

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Make the work visible; public Slack threads drove adoption; private per-run channels did not build team habits

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Playbooks need enablement; workshops and hackathons turn repeated operational work into reusable playbooks

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Deterministic verification before probabilistic judgment; SQL linting and EXPLAIN before deeper validation; LLM-as-judge + DeepEval

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for engineering task → reviewed agent output; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The platform spans several workflows; the cited engineering examples retain human review of agent output.
Observation date
2025-11-11

Sources

  1. Delegating Engineering Work To Cloud-Based Agents (Flux) · Preserved Markdownhttps://x.com/AIatDoorDash/status/2087285008906240193social-post · direct-participant · Last source verification: 2026-08-31
  2. Beyond single agents: DoorDash's collaborative AI ecosystem · Preserved Markdownhttps://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/engineering-blog · first-party · Last source verification: 2026-08-31
  3. Delegating Engineering Work To Cloud-Based Agents · Preserved Markdownhttps://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
DropboxPlatform

Nova

An internal platform for coding agents: engineers launch parallel sessions and internal systems invoke agents inside automated SDLC workflows.

CodingCI triageOn-callMaintenance
agent-assisted SDLC workflow → accepted change: Unknown
Operating model, claims & sources

Scoped operating models

agent-assisted SDLC workflow → accepted changeUnknown · Level unknown

Summary and context

Reported metrics

Headline claim

Dozens of agents can run in parallel from one runbook

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Dropbox
Scope
Migration-owner orchestration of dozens of agents from a shared runbook; qualitative capacity description
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Flaky-test remediation (Deflaker): 100+ validation runs

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Dropbox
Scope
CI validation runs per proposed Deflaker flaky-test fix
Denominator
Unknown
Method
Run the test 100 or more times depending on its failure rate; retry capped at five fix attempts
Observation date
Unknown
Key observation

Predecessor Goose-based migrator used across thousands of migration entries before workflows moved onto Nova

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Dropbox
Scope
Predecessor Goose-based migrator, before workflows moved onto Nova
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Dozens of agents launchable from one runbook

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Dropbox
Scope
Migration-owner orchestration of dozens of agents from a shared runbook; qualitative capacity description
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Harness

Validation loop (propose → validate → feed back) with continue_on_validation_failure and max_iterations (~5); branch management kept outside the agent

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Operating model evidence

Operating model assessment

Unclassified for agent-assisted SDLC workflow → accepted change; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The source documents human participation but does not locate one consistent attention boundary across Nova workflows.
Observation date
2026-05-22

Sources

  1. Introducing Nova, our internal platform for coding agents · Preserved Markdownhttps://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agentsengineering-blog · first-party · Last source verification: 2026-08-31
  2. Hacker News submission for Nova · Preserved Markdownhttps://news.ycombinator.com/item?id=48235065hn-thread · community · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
FlexTask agent

AI Investigation Agent

A Slack agent for HSA/FSA payment operations that traces a payment end-to-end and, when it finds a software bug, prepares a PR with a proposed fix.

Finance opsOn-callCoding
payment investigation → proposed code fix: Work product review
Operating model, claims & sources

Scoped operating models

payment investigation → proposed code fixWork product review · Level 3

Summary and context

Summary

A Slack agent for HSA/FSA payment operations that traces a payment end-to-end and, when it finds a software bug, prepares a PR with a proposed fix.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Lessons and interpretation

Lesson

Start where correctness is observable; payment investigation produces artifacts (a trace, a hypothesis, a diff) that can be checked

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

An ops agent that can prepare a fix (not just a report) closes the loop from investigation to code

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for payment investigation → proposed code fix; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes a trace, diagnosis, and proposed pull request returned for review.
Observation date
2026

Sources

  1. The Flex AI Investigation Agent for HSA/FSA payments · Preserved Markdownhttps://www.withflex.com/blog/the-flex-ai-investigation-agent-for-hsa-fsa-paymentsengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
GitHubTask agent

Qubot

GitHub's internal data-analytics agent, powered by GitHub Copilot. Any GitHub employee can ask a question about the company data warehouse in plain language and get an answer within seconds.

Data
data question → warehouse answer: Unknown
Operating model, claims & sources

Scoped operating models

data question → warehouse answerUnknown · Level unknown

Summary and context

Summary

GitHub's internal data-analytics agent, powered by GitHub Copilot. Any GitHub employee can ask a question about the company data warehouse in plain language and get an answer within seconds.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Hundreds of users run thousands of queries; data questions in internal Slack channels dropped

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
GitHub reported the adoption figures in its own engineering blog without independent verification.
Reported by
GitHub
Scope
Internal GitHub Qubot users and queries; no exact count or measurement window
Denominator
Unknown
Method
Unknown
Observation date
2026-06
Key observation

Hundreds of users run thousands of queries

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
GitHub reported the adoption figures in its own engineering blog.
Reported by
GitHub
Scope
Internal GitHub Qubot users and queries; no exact count or measurement window
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Other reported details and interpretation

Key observation

Volume of data questions in internal data and analytics Slack channels dropped

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Company reports a qualitative decrease without before-and-after counts.
Scope
Questions in GitHub internal data and analytics Slack channels

Operating model evidence

Operating model assessment

Unclassified for data question → warehouse answer; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The public evidence does not document where human attention returns in the question-and-answer flow.
Observation date
2026-06

Sources

  1. How we built an internal data analytics agent · Preserved Markdownhttps://github.blog/ai-and-ml/github-copilot/how-we-built-an-internal-data-analytics-agent/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
HarveyPlatform

Spectre

Harvey's internal collaborative cloud agent platform; reacts to incidents, bug reports, and Slack messages and produces reviewable diffs, branches, and PRs.

CodingCode reviewOn-callSecurity
incident or request → reviewable diff or pull request: Work product review
Operating model, claims & sources

Scoped operating models

incident or request → reviewable diff or pull requestWork product review · Level 3

Summary and context

Architecture and primitives

Lessons and interpretation

Lesson

Harvey keeps its product-agent and security-agent platforms on separate substrates because they have different trust boundaries

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for incident or request → reviewable diff or pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source explicitly frames diffs, branches, and pull requests as reviewable outputs.
Observation date
2026

Sources

  1. Building Spectre; internal collaborative cloud agent platform · Preserved Markdownhttps://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platformengineering-blog · first-party · Last source verification: 2026-08-31
  2. Building an agentic security operations center · Preserved Markdownhttps://www.harvey.ai/blog/building-an-agentic-security-operations-centerengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-08-31
HubSpotTask agent

Sidekick

HubSpot's internal AI code-review agent. Sidekick reviews every pull request and uses a multi-model Judge Agent to filter comments before posting. Its review implementation moved from Claude Code on Crucible Kubernetes workloads to Aviator, HubSpot's internal Java agent framework; the later report does not specify Aviator's execution isolation.

Code review
pull request → AI review comments: Work product review
Operating model, claims & sources

Scoped operating models

pull request → AI review commentsWork product review · Level 3

Summary and context

Summary

HubSpot's internal AI code-review agent. Sidekick reviews every pull request and uses a multi-model Judge Agent to filter comments before posting. Its review implementation moved from Claude Code on Crucible Kubernetes workloads to Aviator, HubSpot's internal Java agent framework; the later report does not specify Aviator's execution isolation.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Reviews every pull request and cut engineer feedback time by 90%

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
HubSpot reported the figures in its own engineering blog without independent verification.
Reported by
HubSpot
Scope
Time for engineers to receive code feedback from Sidekick; not overall PR completion time
Denominator
Unknown
Method
Unknown
Observation date
2026-03
Key observation

Reviews every pull request

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
HubSpot reported the figures in its own engineering blog.
Reported by
HubSpot
Scope
Pull-request coverage after the six-month rollout
Denominator
HubSpot pull requests
Method
Unknown
Observation date
Unknown
Key observation

Engineer feedback time cut by 90%

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
HubSpot reported the figures in its own engineering blog.
Reported by
HubSpot
Scope
Time for engineers to receive code feedback from Sidekick; not overall PR completion time
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Over 80% thumbs-up reaction rate on review feedback during the preceding couple of months

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
HubSpot reported the figures in its own engineering blog.
Reported by
HubSpot
Scope
Developer emoji reactions on review comments during the preceding couple of months
Denominator
Thumbs-up and thumbs-down reactions; not all developers or all reviews
Method
Emoji reactions and replies on review comments
Observation date
Unknown

Architecture and primitives

Harness

Aviator, an internal Java agent framework; replaced the earlier Claude Code review implementation on Crucible

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Sandbox

unknown

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
Current review runs on Aviator; its execution isolation is not specified. Crucible Kubernetes workloads describe the predecessor implementation.

Operating model evidence

Operating model assessment

Level 3 for pull request → AI review comments; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The engineering blog describes engineers acting on Sidekick review comments, which locates human attention at work-product review.
Observation date
2026-03

Sources

  1. Automated code review, the 6-month evolution · Preserved Markdownhttps://product.hubspot.com/blog/automated-code-review-the-6-month-evolutionengineering-blog · first-party · Last source verification: 2026-08-31
  2. Cloud coding agents at HubSpot · Preserved Markdownhttps://product.hubspot.com/blog/cloud-coding-agents-at-hubspotengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
LinearTask agent

Linear Agent

A native agent that synthesizes workspace context, triages, creates follow-up work, runs coding sessions, and executes scheduled/event-driven 'Loops'; used by Linear's own CX, Product, and Engineering teams.

SupportCustomer successCoding
assigned coding work → agent-created change: Work product review
Operating model, claims & sources

Scoped operating models

assigned coding work → agent-created changeWork product review · Level 3

Summary and context

Summary

A native agent that synthesizes workspace context, triages, creates follow-up work, runs coding sessions, and executes scheduled/event-driven 'Loops'; used by Linear's own CX, Product, and Engineering teams.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Sandbox

unknown

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The preserved sources do not document an execution sandbox; unknown does not mean absent.
Model

Codex was used for internal pull-request review; other model choices are not detailed in the preserved sources

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Triage Intelligence (auto-route, dedup, label); GitHub; testing Code Intelligence + custom MCP servers

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Semantic/vector search evolved into agentic context acquisition across the workspace; Datadog/Sentry customer context

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

The Agent SDK gives agents explicit identities, scoped OAuth tokens, assignable/mentionable handles, and visible human delegation; issues stay assigned to a human; 'an agent cannot be held accountable'

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Auto-routes issues, flags duplicates, suggests labels

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Scheduled or event-driven agent runs that execute recurring work

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Keep the agent close to the source of work (Intercom, Slack, Linear); the best workflows live where work already happens

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Gradual autonomy; start by asking for suggestions, observe, add guidance, only automate once proven reliable

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Break work into small steps to keep coding agents focused and successful

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

One Linear engineer reports that agent mistakes reveal possible failure modes during review

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Close the loop; auto-notify the customer when their request ships

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for assigned coding work → agent-created change; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source documents delegated coding output while a human remains accountable for the assigned issue.
Observation date
2026-08-11

Sources

  1. How we built Linear Agent · Preserved Markdownhttps://linear.app/now/how-we-built-linear-agentengineering-blog · first-party · Last source verification: 2026-08-31
  2. How we use Linear Agent at Linear · Preserved Markdownhttps://linear.app/now/how-we-use-linear-agent-at-linearengineering-blog · first-party · Last source verification: 2026-08-31
  3. Our approach to building the Agent Interaction SDK · Preserved Markdownhttps://linear.app/now/our-approach-to-building-the-agent-interaction-sdkengineering-blog · first-party · Last source verification: 2026-08-31
  4. Hacker News submission for how Linear built its agent · Preserved Markdownhttps://news.ycombinator.com/item?id=49252304hn-thread · community · Last source verification: 2026-08-31
  5. Hacker News submission for the Linear Agent public beta · Preserved Markdownhttps://news.ycombinator.com/item?id=48503334hn-thread · community · Last source verification: 2026-08-31
  6. Introducing Linear Agent · Preserved Markdownhttps://linear.app/changelog/2026-03-24-introducing-linear-agentrelease · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
MicrosoftBackground agent

PRAssistant

Microsoft's internal AI code-review agent, built by the Developer Division Data and AI team. When an engineer creates a pull request, PRAssistant joins as a reviewer and leaves comments like a human reviewer. It is a distinct internal build that predates and later informed GitHub Copilot Pull Request Reviews.

Code review
pull request → AI review comments: Work product review
Operating model, claims & sources

Scoped operating models

pull request → AI review commentsWork product review · Level 3

Summary and context

Summary

Microsoft's internal AI code-review agent, built by the Developer Division Data and AI team. When an engineer creates a pull request, PRAssistant joins as a reviewer and leaves comments like a human reviewer. It is a distinct internal build that predates and later informed GitHub Copilot Pull Request Reviews.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Supports more than 90% of Microsoft PRs, impacting over 600,000 pull requests per month

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Microsoft reported the figures in its own engineering blog without independent verification.
Reported by
Microsoft
Scope
PRs supported by the internal AI review assistant across Microsoft
Denominator
Company pull requests for coverage share
Method
Unknown
Observation date
2025-07
Key observation

More than 90% of pull requests across the company

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Microsoft reported the figures in its own engineering blog.
Reported by
Microsoft
Scope
PRs supported by the internal AI review assistant across Microsoft
Denominator
Company pull requests for coverage share
Method
Unknown
Observation date
Unknown
Key observation

More than 600,000 pull requests impacted per month

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Microsoft reported the figures in its own engineering blog.
Reported by
Microsoft
Scope
PRs supported by the internal AI review assistant across Microsoft
Denominator
Company pull requests for coverage share
Method
Unknown
Observation date
Unknown
Key observation

About 5,000 repositories in early onboarding

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Microsoft reported the figures in its own engineering blog.
Reported by
Microsoft
Scope
Repositories onboarded in early AI code-review experiments
Denominator
Unknown
Method
Early experiments and data science studies across onboarded repositories
Observation date
Unknown

Architecture and primitives

Operating model evidence

Operating model assessment

Level 3 for pull request → AI review comments; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The engineering blog describes engineers acting on PRAssistant review comments, which locates human attention at work-product review.
Observation date
2025-07

Sources

  1. Enhancing code quality at scale with AI-powered code reviews · Preserved Markdownhttps://devblogs.microsoft.com/engineering-at-microsoft/enhancing-code-quality-at-scale-with-ai-powered-code-reviews/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
monday.comAgent system

Sphera / Atlas / Morphex

An internal agent system on Amazon Bedrock where agents have identities, managers, scopes, and performance scores; Atlas ships features, Morphex ships PRs autonomously.

CodingCode review
Atlas or Morphex feature task → tested and merged pull request: Outcome review
Operating model, claims & sources

Scoped operating models

Atlas or Morphex feature task → tested and merged pull requestOutcome review · Level 4

Summary and context

Reported metrics

Headline claim

Morphex: 19 of 20 PRs merge without human review

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
monday.com
Scope
Morphex PRs that merge automatically after CI and Guardrails pass
Denominator
Morphex pull requests; sample size and period not supplied
Method
Unknown
Observation date
Unknown
Key observation

Morphex: 19 of 20 PRs merge automatically without human review

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
monday.com
Scope
Morphex PRs that merge automatically after CI and Guardrails pass
Denominator
Morphex pull requests; sample size and period not supplied
Method
Unknown
Observation date
Unknown
Key observation

90% of Builders use AI coding tools monthly; adoption nearly doubled year over year

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
monday.com
Scope
Monthly AI coding-tool adoption among monday Builders, including engineers, PMs, analysts, and designers
Denominator
Builders; not exclusively engineers
Method
Unknown
Observation date
Unknown
Key observation

Per-engineer PR throughput increased by more than 50%

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
monday.com
Scope
Per-engineer pull-request throughput
Denominator
Engineers; baseline period and cohort size not specified
Method
Unknown
Observation date
Unknown
Key observation

Guardrails catches ~25% of agent PRs before human review; low single-digit revert rate

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
monday.com
Scope
Recent cut of top PR-generating agents; Guardrails rejection before review and reverts among merged PRs
Denominator
Agent PRs for Guardrails rejection; merged PRs for revert rate
Method
Unknown
Observation date
Unknown

Architecture and primitives

Lessons and interpretation

Operating model evidence

Operating model assessment

Level 4 for Atlas or Morphex feature task → tested and merged pull request; human attention boundary: outcome-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The secondary source reports predominantly automatic merges and automated guardrails, while humans manage tasks and outcomes.
Observation date
2026

Sources

  1. AI Teammates: how monday.com runs production AI agents on Amazon Bedrock · Preserved Markdownhttps://aws.amazon.com/blogs/machine-learning/ai-teammates-how-monday-com-runs-production-ai-agents-on-amazon-bedrock/case-study · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
NotionPlatform

Custom Agents

Notion's Custom Agents platform, dogfooded internally across non-engineering teams such as IT ticketing, supply chain, procurement, and recruiting. By the end of alpha testing, Notion had more than 3,000 internal Custom Agents. Notion's own security team is one of the most active internal users. Notion rebuilt the agent harness three to five times as frontier models improved.

SupportFinance opsRecruitmentSecurity
cross-team internal tasks → Custom Agents output: Unknown
Operating model, claims & sources

Scoped operating models

cross-team internal tasks → Custom Agents outputUnknown · Level unknown

Summary and context

Summary

Notion's Custom Agents platform, dogfooded internally across non-engineering teams such as IT ticketing, supply chain, procurement, and recruiting. By the end of alpha testing, Notion had more than 3,000 internal Custom Agents. Notion's own security team is one of the most active internal users. Notion rebuilt the agent harness three to five times as frontier models improved.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

More than 3,000 internal Custom Agents by end of alpha testing

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Notion reported the agent count in its own engineering blog without independent verification.
Reported by
Notion
Scope
Internal Notion Custom Agents at the end of alpha; excludes the separate customer alpha count
Denominator
Unknown
Method
Unknown
Observation date
2026-04
Key observation

More than 3,000 internal Custom Agents by end of alpha testing

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Notion reported the agent count in its own engineering blog.
Reported by
Notion
Scope
Internal Notion Custom Agents at the end of alpha; excludes the separate customer alpha count
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Agent harness rebuilt three to five times as models improved

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Participants give differing approximate rebuild counts for harness, framework, and feature; not a precise engineering inventory.
Reported by
Notion
Scope
Notion agent/framework rebuilds recalled by participants; estimates range from three to five
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Other reported details and interpretation

Key observation

Notion's security team is one of the most active internal users

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Qualitative first-party description; no comparative activity count is supplied.
Scope
Notion security team internal Custom Agent use

Operating model evidence

Operating model assessment

Unclassified for cross-team internal tasks → Custom Agents output; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
Custom Agents spans many teams and workflows, so no single human-attention boundary applies.
Observation date
2026-04

Sources

  1. Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space) · Preserved Markdownhttps://latent.space/p/notionpodcast · direct-participant · Last source verification: 2026-08-31
  2. How we built security into Custom Agents · Preserved Markdownhttps://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agentsengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
PlaidTask agent

AI Annotator

Plaid's internal labeling agent for its own model training. AI Annotator automates large-scale labeling of anonymized transaction data, with human oversight on the labeled output. Plaid reports greater than 95% human alignment at a lower cost and time than manual labeling.

Data
raw transactions → labeled training data: Work product review
Operating model, claims & sources

Scoped operating models

raw transactions → labeled training dataWork product review · Level 3

Summary and context

Summary

Plaid's internal labeling agent for its own model training. AI Annotator automates large-scale labeling of anonymized transaction data, with human oversight on the labeled output. Plaid reports greater than 95% human alignment at a lower cost and time than manual labeling.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Greater than 95% human alignment at lower cost and time than manual labeling

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figure in its own blog without independent verification.
Reported by
Plaid
Scope
Transaction labels generated by AI Annotator in early use
Denominator
Labels compared with human judgments; sample size unspecified
Method
Unknown
Observation date
2025-06
Key observation

Greater than 95% human alignment with labeled data

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figure in its own blog.
Reported by
Plaid
Scope
Transaction labels generated by AI Annotator in early use
Denominator
Labels compared with human judgments; sample size unspecified
Method
Unknown
Observation date
Unknown
Key observation

Lower cost and time than manual labeling

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Qualitative company comparison; neither actual costs nor elapsed-time measurements are supplied.
Reported by
Plaid
Scope
AI transaction annotation cost and time relative to manual labeling
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Knowledge

Anonymized Plaid transaction data

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Operating model evidence

Operating model assessment

Level 3 for raw transactions → labeled training data; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The reported greater than 95% human alignment implies that people review the labeled output.
Observation date
2025-06

Sources

  1. AI agents at Plaid (June 2025) · Preserved Markdownhttps://plaid.com/blog/ai-agents-june-2025/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
PlaidTask agent

Fix My Connection

Plaid's internal agent for bank-integration reliability. Fix My Connection proactively detects bank-integration failures and generates repair scripts automatically. Plaid reports more than 2 million successful user-permissioned logins and a 90% reduction in the average time to fix a degradation.

OpsMaintenance
integration degradation → repaired connection: Outcome review
Operating model, claims & sources

Scoped operating models

integration degradation → repaired connectionOutcome review · Level 4

Summary and context

Summary

Plaid's internal agent for bank-integration reliability. Fix My Connection proactively detects bank-integration failures and generates repair scripts automatically. Plaid reports more than 2 million successful user-permissioned logins and a 90% reduction in the average time to fix a degradation.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

More than 2 million successful logins and 90% faster average repair

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figures in its own blog without independent verification.
Reported by
Plaid
Scope
Automated repair of bank connections and the successful user-permissioned logins enabled by those repairs
Denominator
Unknown
Method
Unknown
Observation date
2025-06
Key observation

More than 2 million successful user-permissioned logins

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figure in its own blog.
Reported by
Plaid
Scope
Automated repair of bank connections and the successful user-permissioned logins enabled by those repairs
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Average time to fix a degradation reduced by 90%

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figure in its own blog.
Reported by
Plaid
Scope
Automated repair of bank connections and the successful user-permissioned logins enabled by those repairs
Denominator
Average degradation-repair time before automated repairs; no baseline duration provided
Method
Unknown
Observation date
Unknown

Architecture and primitives

Tool access

Plaid bank-integration infrastructure

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Operating model evidence

Operating model assessment

Level 4 for integration degradation → repaired connection; human attention boundary: outcome-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
Plaid measures success by outcomes such as successful logins rather than per-repair inspection.
Observation date
2025-06

Sources

  1. AI agents at Plaid (June 2025) · Preserved Markdownhttps://plaid.com/blog/ai-agents-june-2025/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
PlaidSupporting pattern

Internal MCP server

Plaid's central internal Model Context Protocol server. Plaid built it because third-party MCP servers could not reach its internal data. The server integrates more than 20 tools and several internal services such as Jira, application logs, and data schemas, behind Plaid's identity-aware proxy and centralized authorization. Plaid reports thousands of tool calls and dozens of agents built on the server. Separately, Claude Code and Cursor are used by more than 80% of Plaid engineers; server adoption is not quantified.

Coding
engineer request → internal tool access: Unknown
Operating model, claims & sources

Scoped operating models

engineer request → internal tool accessUnknown · Level unknown

Summary and context

Summary

Plaid's central internal Model Context Protocol server. Plaid built it because third-party MCP servers could not reach its internal data. The server integrates more than 20 tools and several internal services such as Jira, application logs, and data schemas, behind Plaid's identity-aware proxy and centralized authorization. Plaid reports thousands of tool calls and dozens of agents built on the server. Separately, Claude Code and Cursor are used by more than 80% of Plaid engineers; server adoption is not quantified.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Dozens of agents rely on the internal MCP server; Claude Code and Cursor are used by over 80% of engineers

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figures in its own engineering blog without independent verification.
Reported by
Plaid
Scope
Claude Code and Cursor adoption among Plaid engineers; separate from internal MCP server adoption
Denominator
Plaid engineers for the AI-client usage share
Method
Unknown
Observation date
2025
Key observation

Claude Code and Cursor are used by over 80% of Plaid engineers; the source does not report internal MCP server adoption share

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figures in its own engineering blog.
Reported by
Plaid
Scope
Claude Code and Cursor adoption among Plaid engineers; separate from internal MCP server adoption
Denominator
Plaid engineers for the AI-client usage share
Method
Unknown
Observation date
Unknown
Key observation

Thousands of tool calls and dozens of agents built on it

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Plaid reported the figures in its own engineering blog.
Reported by
Plaid
Scope
Tool calls and agents relying on the internal MCP server across engineering, product, and support
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Harness

Central internal MCP server that fronts vendor AI clients

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

More than 20 tools and several internal services (Jira, logs, schemas)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Behind Plaid's identity-aware proxy and centralized authorization

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Operating model evidence

Operating model assessment

Unclassified for engineer request → internal tool access; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The MCP server is a tool-access layer, not a workflow with a single human-attention boundary.
Observation date
2025

Sources

  1. The Plaid internal MCP server · Preserved Markdownhttps://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdbengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
PostHogBackground agent

StampHog

A GitHub-label-triggered PR approval agent that applies fail-closed deterministic safety gates, asks an LLM to check for showstoppers, autonomously approves eligible changes, and refuses or escalates the rest.

Code review
eligible pull request → approval decision: Exception only
Operating model, claims & sources

Scoped operating models

eligible pull request → approval decisionException only · Level 5

Summary and context

Summary

A GitHub-label-triggered PR approval agent that applies fail-closed deterministic safety gates, asks an LLM to check for showstoppers, autonomously approves eligible changes, and refuses or escalates the rest.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Handled 1,600 PRs in the previous month, as reported on July 9, 2026; roughly one in three merged main-repository PRs received its final approval during the reported quarter

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
PostHog reports two different windows in its July 9, 2026 article; exact monthly and quarterly boundaries are not supplied.
Reported by
PostHog
Scope
Monthly PRs handled autonomously and quarterly final approvals in the main repository
Denominator
Merged main-repository PRs for the quarterly share; absolute handled PRs for the monthly count
Method
Company-reported production usage
Observation date
2026-07
Key observation

1,600 PRs handled autonomously in the previous month, as reported on July 9, 2026

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
July 9 is the report date; the source says last month without exact measurement boundaries.
Reported by
PostHog
Scope
PRs handled autonomously in the previous month, as reported on July 9, 2026; exact boundaries unspecified
Denominator
Unknown
Method
Company-reported production usage
Observation date
2026-07-09
Key observation

Roughly one in three PRs merged into PostHog's main repository received StampHog's final approval during the reported quarter

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
PostHog
Scope
Main-repository merged PRs receiving StampHog final approval during the reported quarter
Denominator
Merged PRs in the main repository
Method
Company-reported production usage
Observation date
2026-07
Key observation

20% of PRs approved by StampHog in the July 28, 2026 report, at approximately $300 per month in tokens

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
PostHog
Scope
PostHog PRs approved by StampHog and monthly token cost in the July 28, 2026 report
Denominator
PostHog PRs in the report's scope
Method
Company-reported production usage and token spend
Observation date
2026-07

Architecture and primitives

Harness

A GitHub Action invokes a Python pipeline that fetches and classifies a PR, applies hard gates, waits for in-flight reviewer bots, runs an LLM review, and posts a verdict

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Model

Claude through the Claude Agent SDK, with Read, Grep, and Glob tools

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Reads the diff and repository files plus trusted review-state, discussion, ownership, and reviewer signals

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Repository-specific deny categories, size and risk tiers calibrated from prior human approvals, review guidance, and ownership data

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

A dedicated Anthropic organization secret and a StampHog GitHub App token whose approvals satisfy branch protection

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Each run emits a versioned JSON evidence bundle retained as a CI artifact for 30 days; a sticky GitHub comment carries non-approval verdict history and labels preserve retry state

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Draft state, conflicts, requested changes, sensitive paths, size ceilings, and risk tiers can block AI approval; the LLM may tighten but never loosen a gate

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Eligible changes can be approved while risky, ambiguous, or insufficiently assured changes are refused or escalated to a suitable human reviewer

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Each run records PR metadata, classification, gate results, reviewer output, and the final verdict

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Use deterministic controls for known risks and allow the LLM to make approval stricter, never more permissive

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source code explicitly implements and documents this safety invariant.
Lesson

Calibrate thresholds and deny categories from repository history rather than treating small diffs as inherently safe

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source code documents calibration against historical approval outcomes.
Lesson

Fail closed and preserve retry state when dependencies, credentials, or concurrent reviewer bots are unavailable

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source code explicitly documents fail-closed and retry behavior.

Operating model evidence

Operating model assessment

Level 5 for eligible pull request → approval decision; human attention boundary: exception-only.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
Eligible pull requests are approved automatically; risky or ambiguous cases are refused or routed to a human.
Observation date
2026-07-09

Sources

  1. Stop being the code review bottleneck · Preserved Markdownhttps://posthog.com/newsletter/code-review-tipscorporate-article · first-party · Last source verification: 2026-08-31
  2. 10,000 PRs a month is easy: How devex is evolving at PostHog · Preserved Markdownhttps://posthog.com/blog/10k-prs-a-monthengineering-blog · first-party · Last source verification: 2026-08-31
  3. StampHog PR approval agent source code and documentation · Preserved Markdownhttps://github.com/PostHog/posthog/tree/988c9031bb93c74bafcdfb670c01497c79a4f644/tools/pr-approval-agentsource-code · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
RampBackground agent

Inspect

A background coding agent that closes the loop on verifying its own work; runs tests, reviews telemetry, queries feature flags, visually verifies the frontend; now also monitoring production and proposing fixes; also a platform that hosts many internal agents.

CodingCode reviewOn-call
Inspect coding task → reviewed production merge: Work product review
Operating model, claims & sources

Scoped operating models

Inspect coding task → reviewed production mergeWork product review · Level 3

Summary and context

Summary

A background coding agent that closes the loop on verifying its own work; runs tests, reviews telemetry, queries feature flags, visually verifies the frontend; now also monitoring production and proposing fixes; also a platform that hosts many internal agents.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

75% of Ramp's merged PRs raised by Inspect sessions (May 2026)

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Ramp merged PRs raised by Inspect sessions by May 2026
Denominator
Merged Ramp pull requests
Method
Unknown
Observation date
2026-05
Key observation

Around 60% of Ramp PRs authored by Inspect by January 2026

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Ramp PRs authored by Inspect by January 2026
Denominator
Ramp pull requests in the source adoption history
Method
Unknown
Observation date
2026-01
Key observation

75% of merged PRs raised by Inspect sessions (May 2026)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Ramp
Scope
Ramp merged PRs raised by Inspect sessions by May 2026
Denominator
Merged Ramp pull requests
Method
Unknown
Observation date
2026-05
Key observation

Around 30% of merged frontend and backend PRs in the earlier first-party report, after a couple of months of adoption

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Ramp
Scope
Merged PRs in Ramp frontend and backend repositories after the first couple of months of adoption
Denominator
Merged PRs in frontend and backend repositories
Method
Unknown
Observation date
Unknown
Key observation

Around 90% of PRs merged into the Inspect repository come from Inspect sessions in the later interview report

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
PRs merged into the Inspect repository
Denominator
Merged PRs in the Inspect repository
Method
Unknown
Observation date
Unknown
Key observation

One million total Inspect sessions crossed in July 2026

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Cumulative Inspect sessions crossing the million mark in July 2026
Denominator
Unknown
Method
Unknown
Observation date
2026-07
Key observation

Under 5 seconds to spin up a fully provisioned remote dev environment

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Provisioning a fully configured remote development environment
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

150+ engineers contributed to the Inspect codebase in the later interview report

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Engineers who contributed to the Inspect codebase in the later interview report
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

5.5-person Inspect team (four engineers, a director, and a part-time PM)

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Inspect team staffing: four engineers, one director, and a part-time PM
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

More than 80% of Inspect is written in Inspect sessions

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Reported by
Ramp
Scope
Inspect code written in Inspect sessions; distinct from merged PR share
Denominator
Inspect code; the source does not define a code-volume counting method
Method
Unknown
Observation date
Unknown

Architecture and primitives

Sandbox

Modal sandboxes; per-repo images rebuilt every 30 min from snapshots; warm-on-keystroke; a pool of warm sandboxes

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

OpenCode (server-first) as the agent runtime; a plugin blocks writes until sync completes; expanded into production monitoring and self-maintenance; Inspect itself is built with React/Vite, Cloudflare Durable Objects, SQLite, and the Cloudflare Agents SDK

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Wired into Sentry, Datadog, LaunchDarkly, Braintrust, GitHub, Slack, Buildkite; monitors production, triages issues, proposes fixes; debugging queries a sanitized read-only production DB replica and Snowflake

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Skills that encode how Ramp ships; repo images with the full dev env (Vite, Postgres, Redis, RabbitMQ, Temporal, Chromium, VS Code Server)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

GitHub auth per user; the sandbox pushes the branch, an API opens the PR with the user's token (no self-approval); production merges retain human review

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Other reported details and interpretation

Key observation

Inspect underpins internal agents including ReviewBuddy, Oncall Assistant, Testo, Ramp Research, Voice of the Customer, and error automations

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked participant or independent source reports the claim.
Scope
Examples of internal agents built on Inspect
Key observation

Design goal: session speed should be limited only by model-provider time-to-first-token

Opinion · Reported · Medium confidence
Evidence and qualifications
Confidence reason
The source says session speed should only be limited by model time-to-first-token; this is a design goal, not a measured result.
Scope
Target session startup speed, excluding precompleted cloning and installation

Lessons and interpretation

Lesson

Own the tooling; it only has to work on your code, which lets you build something more powerful than off-the-shelf

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Work in public spaces to create virality loops; let the product do the talking, don't mandate

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Ramp argues that a fast background agent can add remote resources and concurrency to the same model

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Move as much as possible into the image-build step so users never wait on setup

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

The v1 Chrome extension saw little adoption; the pivot to a centrally configured remote dev environment with a coding agent on top drove adoption

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source reports that v1 saw little adoption and that the November 2025 pivot to a remote dev environment preceded rapid adoption.

Operating model evidence

Operating model assessment

Level 3 for Inspect coding task → reviewed production merge; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source explicitly states that Inspect cannot self-approve and production merges retain human review.
Observation date
2026-05

Sources

  1. Why We Built Our Own Background Agent (Inspect) · Preserved Markdownhttps://engineering.ramp.com/post/why-we-built-our-background-agentengineering-blog · first-party · Last source verification: 2026-08-31
  2. Ramp x Linear (75% of merged PRs via Inspect) · Preserved Markdownhttps://linear.app/customers/rampcase-study · first-party · Last source verification: 2026-08-31
  3. Hacker News project discussion inspired by Ramp Inspect · Preserved Markdownhttps://news.ycombinator.com/item?id=48042123hn-thread · community · Last source verification: 2026-08-31
  4. Why We Built Our Own Background Agent (Inspect) (Ramp Builders URL) · Preserved Markdownhttps://builders.ramp.com/post/why-we-built-our-background-agentengineering-blog · first-party · Last source verification: 2026-08-31
  5. Why Ramp built its own in-house coding agent, Inspect · Preserved Markdownhttps://newsletter.pragmaticengineer.com/p/why-ramp-built-inspectnews · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
ReplitOrchestration system

Manager agent (agent-of-agents)

An internal agent-of-agents stack where every employee gets a manager agent that spawns multiple agents for verifiable work and escalates judgment to humans.

CodingCode reviewSupportResearchData
objective → verifiable multi-agent work product: Work product review
Operating model, claims & sources

Scoped operating models

objective → verifiable multi-agent work productWork product review · Level 3

Summary and context

Summary

An internal agent-of-agents stack where every employee gets a manager agent that spawns multiple agents for verifiable work and escalates judgment to humans.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

2.9x code output for a consistent author cohort; review latency, PR reversions, and incident trends reported flat

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Replit
Scope
Code output for a consistent author cohort across early January to late June; separate from company-wide hiring effects
Denominator
Same cohort of authors before and after
Method
Comparison of contributed code for a consistent author cohort; raw company-wide lines of code increased 5.8x
Observation date
Unknown
Key observation

2.9x code output for a consistent author cohort from early January to late June

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Replit
Scope
Code output for a consistent author cohort across early January to late June; separate from company-wide hiring effects
Denominator
Same cohort of authors before and after
Method
Comparison of contributed code for a consistent author cohort; raw company-wide lines of code increased 5.8x
Observation date
Unknown
Key observation

No corresponding deterioration in review/reversion/incident metrics

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Replit
Scope
Company code review latency, PR reversion rates, and incidents opened during increased code output
Denominator
Unknown
Method
Company comparison of review latency, PR reversion rates, and incident trends
Observation date
Unknown

Architecture and primitives

Sandbox

microVMs and remote filesystems behind access policies, token proxies, audit logging, and a ZeroTrust network

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

Fleet/loop orchestration: a manager agent launches parallel agents for verifiable work and escalates judgment

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Model

Not specified

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Investigates incidents, reviews PRs, answers questions, analyzes company data, triages support, researches sales accounts, improves Replit Agent itself

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Manager agent coordinates parallel sub-agents and routes results

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

One human gives an objective; the manager spawns parallel agents for verifiable work and escalates judgment

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Give every employee a manager agent that spawns sub-agents; verifiable work parallelizes, judgment escalates to humans

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Track outcome metrics (reverts, incidents), not activity; output can scale without quality regressions

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for objective → verifiable multi-agent work product; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes autonomous parallel execution followed by human judgment on the resulting work.
Observation date
2026

Sources

  1. The Self-Driving Company · Preserved Markdownhttps://ld.replit.com/blog/self-driving-companyengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
RetoolTask agent

RetoolGPT

Retool's internal assistant, built as a version of ChatGPT with access to Retool's internal Confluence documents, Retool documentation, and Linear tickets. The team deployed it organization-wide in a read-only environment so the whole team could use it.

SupportCoding
internal question → sourced answer: Unknown
Operating model, claims & sources

Scoped operating models

internal question → sourced answerUnknown · Level unknown

Summary and context

Summary

Retool's internal assistant, built as a version of ChatGPT with access to Retool's internal Confluence documents, Retool documentation, and Linear tickets. The team deployed it organization-wide in a read-only environment so the whole team could use it.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Model

Built on ChatGPT

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Retool Confluence documents, Retool documentation, and Linear tickets

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Operating model evidence

Operating model assessment

Unclassified for internal question → sourced answer; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The public evidence does not document where human attention returns in the question-and-answer flow.
Observation date
2025-08

Sources

  1. How we built RetoolGPT · Preserved Markdownhttps://retool.com/blog/how-we-built-retoolgptengineering-blog · first-party · Last source verification: 2026-08-31
  2. AI Build Week, Day 3: How we made RetoolGPT · Preserved Markdownhttps://www.youtube.com/watch?v=8VTdYUBAZsYtalk · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-08-31
SalesforceTask agent

Slackbot

Salesforce was 'customer zero' for the rebuilt Slackbot; an employee agent that finds company context, drafts work, and connects Slack context with Salesforce data; now also an external product.

SupportCustomer successOps
employee request → drafted work: Work product review
Operating model, claims & sources

Scoped operating models

employee request → drafted workWork product review · Level 3

Summary and context

Summary

Salesforce was 'customer zero' for the rebuilt Slackbot; an employee agent that finds company context, drafts work, and connects Slack context with Salesforce data; now also an external product.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Lessons and interpretation

Lesson

Make permission-awareness a construction property, not a prompt instruction; the agent sees only what the employee can see

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for employee request → drafted work; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source describes an employee agent that prepares drafts and contextual work for human use.
Observation date
2026-01-14

Sources

  1. Salesforce announces general availability of Slackbot · Preserved Markdownhttps://www.salesforce.com/ap/news/press-releases/2026/01/14/salesforce-announces-the-general-availability-of-slackbot-your-personal-agent-for-work-sg/release · first-party · Last source verification: 2026-08-31
  2. Interview with Slackbot about connected work context · Preserved Markdownhttps://www.salesforce.com/in/news/stories/salesforce-slackbot-ai-interview/corporate-article · first-party · Last source verification: 2026-08-31
  3. Salesforce rolls out new Slackbot AI agent · Preserved Markdownhttps://venturebeat.com/technology/salesforce-rolls-out-new-slackbot-ai-agent-as-it-battles-microsoft-andnews · independent-secondary · Last source verification: 2026-08-31
  4. Hacker News submission for the Salesforce Slackbot rollout · Preserved Markdownhttps://news.ycombinator.com/item?id=46600760hn-thread · community · Last source verification: 2026-08-31
Entry reviewed 2026-08-31
SentryTask agent

Junior

An open-source Slack agent built at Sentry that acts like an intern; takes tasks, retrieves context across many company systems, and is steered and reviewed by humans. Its CEO argues one general-purpose agent beat several vendor-specific bots.

CodingCode reviewSupportOn-call
assigned task → human-steered and reviewed output: Continuous steering
Operating model, claims & sources

Scoped operating models

assigned task → human-steered and reviewed outputContinuous steering · Level 2

Summary and context

Summary

An open-source Slack agent built at Sentry that acts like an intern; takes tasks, retrieves context across many company systems, and is steered and reviewed by humans. Its CEO argues one general-purpose agent beat several vendor-specific bots.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Open-source (Apache-2.0) Slack agent (~100k lines of TS) used internally at Sentry

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sentry
Scope
Junior codebase size and license in the author report
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Around 100,000 lines of TypeScript excluding tests, evals, docs, and lockfiles

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sentry
Scope
TypeScript lines in Junior excluding tests, evals, documentation, and lockfiles
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

4 months from start to writeup

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sentry
Scope
Author-reported development and iteration time before the writeup
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Sandbox

Vercel serverless functions; Vercel agent-browser sandbox with an on-path proxy for traffic interception; ephemeral containers

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

Custom harness on Pi's SDK; a task broker over Vercel Queues with an inbox -> worker-claim -> interrupt/resume pattern to survive serverless timeouts; 'skills-as-runbooks'

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Model

Claude Sonnet (faster); swappable (Opus as a more expensive option)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Progressive discovery via MCP; by default Junior connects to no provider until the agent requests a tool lookup; plugins connect Sentry, GitHub, Linear, Notion

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Conversation transcripts persisted in Redis; repo search to trace code paths; skill docs (TELEMETRY.md, SOUL.md)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

On-path proxy injection; the model never sees the token because it is not in the sandbox; plugins declare OAuth flows and credential domains; GitHub distinguishes read vs. write

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Incremental transcript updates in Redis; resource subscriptions to GitHub PR events for follow-ups

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

searchMcpTools loads tools on demand instead of dumping every schema into the prompt

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Credentials injected host-side; the sandbox/model never touches a secret

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Survives serverless timeouts via a queue + claim model

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

One general-purpose agent connected to many company systems beats several vendor-specific bots

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Skills-as-runbooks encode operational knowledge the agent can follow

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Stateless compute fights you; serverless functions time out and disappear; model the agent around interrupt/resume

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Unit tests are the wrong yardstick for agents; invest in evals and integration tests instead

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Writes need per-user authorization, not blanket trust

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 2 for assigned task → human-steered and reviewed output; human attention boundary: continuous-steering.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source explicitly describes humans steering and reviewing the agent throughout its work.
Observation date
2026

Sources

  1. Building an Intern (Junior at Sentry) · Preserved Markdownhttps://cra.mr/building-an-intern/engineering-blog · first-party · Last source verification: 2026-08-31
  2. getsentry/junior source repository · Preserved Markdownhttps://github.com/getsentry/juniorrepository · first-party · Last source verification: 2026-08-31
  3. Sentry Labs · Preserved Markdownhttps://labs.sentry.dev/documentation · first-party · Last source verification: 2026-08-31
  4. getsentry/junior Apache 2.0 license · Preserved Markdownhttps://github.com/getsentry/junior/blob/main/LICENSE?plain=1source-code · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
ShopifyPlatform

Aquifer / River

The Aquifer agent platform (session/harness/sandbox split) powers River, a Slack-native coding agent, plus research, migration, and app-security agents.

CodingCode reviewResearchSecurity
River coding request → reviewed pull request: Work product review
Operating model, claims & sources

Scoped operating models

River coding request → reviewed pull requestWork product review · Level 3

Summary and context

Summary

The Aquifer agent platform (session/harness/sandbox split) powers River, a Slack-native coding agent, plus research, migration, and app-security agents.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

1 in 8 merged PRs company-wide coauthored by River

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Shopify
Scope
Merged River-coauthored PRs across Shopify
Denominator
All Shopify merged pull requests
Method
Unknown
Observation date
Unknown
Key observation

River: 59,918 sessions / 30 days across 5,170 Slack channels

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Shopify
Scope
River sessions and distinct Slack channels in a recent 30-day period
Denominator
Unknown
Method
river_sessions domain table, written by River every session
Observation date
Unknown
Key observation

3,536 River-coauthored PRs merged; 1 in 8 merged PRs company-wide

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Shopify
Scope
River-coauthored merged PRs in the recent 30-day period; company-wide PR share
Denominator
All Shopify merged pull requests for the one-in-eight share
Method
river_sessions domain table for reported session-linked counts
Observation date
Unknown
Key observation

Median session 19 min; median 50 tool calls/session

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Shopify
Scope
Median River session duration and tool calls per session
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Sandbox

An execution environment (filesystem, shell, repo, build/test) separated from the harness; Shopify credits the brain-and-hands framing to Anthropic

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

Session (durable; Postgres append-only event log) + Harness (cheap agent loop) + Cell (ephemeral Go runtime); cells die, sessions persist

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Model

The design lets Shopify change the model without changing the sandbox

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Interfaces

slack, github

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Repo + tests + data warehouse + production traces + PR creation; a gateway credentials proxy

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Monorepo 'World' (code + skills + conventions + intent docs + runbooks + AGENTS.md); Nix reproducible envs; Slack-transcript corpus mining

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Gateway credentials proxy; per-profile sandbox policies; Shopify SSO

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Skills loaded on-demand as files, updatable per session; session survival across cell/sandbox/machine death

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Durable identity, disposable loop, isolated execution; swap any layer independently

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Slack-native coding agent in public channels only; visibility drives adoption

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Written-down knowledge mined from successful patterns and public transcripts

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Agent-friendly is human-friendly; monorepo, reproducible envs, written skills, and fast CI help both

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Local agents have a ceiling; private windows mean only the person at the keyboard learns anything

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Session survival is critical; cells die, sandboxes die, machines die; the conversation doesn't

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Treat agents as profiles, not platforms; a new agent is a new bundle on the same substrate

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for River coding request → reviewed pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source documents pull-request creation but does not establish outcome-only supervision.
Observation date
2026

Sources

  1. Under the River · Preserved Markdownhttps://shopify.engineering/under-the-riverengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
SierraTask agent

Pinecone

One company-wide agent that collapsed separate support, analytics, engineering, and sales agents into a single runtime with an MCP Gateway to 45 systems.

CodingCode reviewSupportResearchData
employee request → reviewed agent output: Work product review
Operating model, claims & sources

Scoped operating models

employee request → reviewed agent outputWork product review · Level 3

Summary and context

Reported metrics

Headline claim

More than 75,000 sessions created by 600 people in the month preceding the report

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sierra
Scope
Pinecone users and sessions in the month preceding the report; calendar month unspecified
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

More than 75,000 sessions created by 600 people in the month preceding the report

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sierra
Scope
Pinecone users and sessions in the month preceding the report; calendar month unspecified
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

70% of company PRs opened through Pinecone in the month preceding the report

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Sierra
Scope
Company PRs opened through Pinecone in the month preceding the report
Denominator
Sierra PRs opened in that month; not merged PRs
Method
Unknown
Observation date
Unknown

Architecture and primitives

Sandbox

Agency layer reconciles recoverable Kubernetes runners; conversation/events/checkpoints stay durable separately

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

MCP Gateway connected to 45 systems; a network proxy decides whether privileged requests proceed and injects credentials after approval

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

One gateway spanning 45 systems, enforcing employee permissions at call time

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Collapse departmental bots into one agent; cross-functional jobs don't respect org-chart boundaries

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Session counts and tool calls are usage, not value; track business outcomes instead

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for employee request → reviewed agent output; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The source documents delegated work through a company-wide agent without establishing outcome-only supervision.
Observation date
2026

Sources

  1. Pinecone: harnessing the wisdom of the workforce · Preserved Markdownhttps://sierra.ai/jp/blog/pinecone-harnessing-the-wisdom-of-the-workforceengineering-blog · first-party · Last source verification: 2026-08-31
  2. Building Sierra's MCP Gateway · Preserved Markdownhttps://sierra.ai/blog/building-sierras-mcp-gateway-an-engineering-icebergengineering-blog · first-party · Last source verification: 2026-08-31
  3. Agency: secure, scalable sandboxes for agents · Preserved Markdownhttps://sierra.ai/es/blog/agency-secure-scalable-sandboxes-for-agentsengineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
SlackSupporting pattern

Multi-agent context system

A coordinator/dispatcher multi-agent design with structured context channels for long-running investigations spanning hundreds of steps.

Research
long-running investigation → synthesized report: Unknown
Operating model, claims & sources

Scoped operating models

long-running investigation → synthesized reportUnknown · Level unknown

Summary and context

Reported metrics

Headline claim

Context management for security investigations spanning hundreds of inference requests and megabytes of output

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Slack
Scope
Complex security investigations requiring tailored multi-agent context; qualitative workload scale
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Handles multi-agent runs spanning hundreds of requests and megabytes of output

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Slack
Scope
Complex security investigations requiring tailored multi-agent context; qualitative workload scale
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Context management

Three channels; Director's Journal (working memory), Critic's Review (credibility-weighted findings), Critic's Timeline (deduped chronological synthesis)

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Operating model evidence

Operating model assessment

Unclassified for long-running investigation → synthesized report; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The research pattern documents agent coordination and criticism but not the normal human attention boundary.
Observation date
2026

Sources

  1. Managing context in long-running agentic applications · Preserved Markdownhttps://slack.engineering/managing-context-in-long-run-agentic-applications/engineering-blog · first-party · Last source verification: 2026-08-31
  2. How Slack manages context in long-running multi-agent systems · Preserved Markdownhttps://www.infoq.com/news/2026/04/slack-agent-context-management/news · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
SpotifyAgent system

Honk / Xirp

Honk is Spotify's background coding agent (high confidence); Xirp is the workspace/context/session layer around it (medium confidence). Honk runs on a Claude-Agent-SDK harness in Kubernetes, uses trusted CI tools, and combines formatting and linting with LLM-based diff evaluation.

CodingMigrationsCode review
Honk coding task → verified pull request: Work product review
Operating model, claims & sources

Scoped operating models

Honk coding task → verified pull requestWork product review · Level 3

Summary and context

Summary

Honk is Spotify's background coding agent (high confidence); Xirp is the workspace/context/session layer around it (medium confidence). Honk runs on a Claude-Agent-SDK harness in Kubernetes, uses trusted CI tools, and combines formatting and linting with LLM-based diff evaluation.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

1,500+ merged pull requests generated by Honk

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Spotify
Scope
Cumulative merged Honk-generated PRs reported in Part 1; no measurement cutoff in preserved text
Denominator
Unknown
Method
Company-reported merged pull request count
Observation date
Unknown
Key observation

1,500+ merged AI-generated PRs (Honk)

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Spotify
Scope
Cumulative merged Honk-generated PRs reported in Part 1; no measurement cutoff in preserved text
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

60-90% time savings on migrations vs manual

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Spotify
Scope
Selected code migrations using Honk compared with writing the changes manually
Denominator
Manual completion time for the same migration work
Method
Unknown
Observation date
Unknown
Key observation

Around half of Spotify PRs automated by Fleet Management since mid-2024, including deterministic transformations

Metric · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Spotify
Scope
Fleet Management automated pull requests since mid-2024, including deterministic transforms
Denominator
All Spotify pull requests
Method
Unknown
Observation date
2024

Architecture and primitives

Tool access

Limited, deliberate tool surface; trusted CI tools verify changes, while an internal CLI runs formatting and linting through MCP and evaluates diffs with an LLM judge

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

The session/context surface around the agent rather than an autonomous worker itself

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Autonomy without structure fragments into per-engineer configs; shared org context is the multiplier

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 3 for Honk coding task → verified pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source documents automated verification followed by a pull-request workflow that retains human review.
Observation date
2025

Sources

  1. Spotify's Journey with Our Background Coding Agent (Honk) · Preserved Markdownhttps://engineering.atspotify.com/2025/11/spotifys-background-coding-agent-part-1engineering-blog · first-party · Last source verification: 2026-08-31
  2. Code with Claude: coding is no longer the constraint (Honk v2) · Preserved Markdownhttps://engineering.atspotify.com/2026/6/code-with-claude-coding-is-no-longer-the-constraintengineering-blog · first-party · Last source verification: 2026-08-31
  3. Xirp - Powered by Spotify Portal · Preserved Markdownhttps://xirp.spotify.com/documentation · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
StripeBackground agent

Minions

Homegrown one-shot coding agents that read work context, locate the right repo/workspace, implement a change end-to-end, and produce a PR for human review.

CodingCode review
work context → merge-ready pull request: Work product review
Operating model, claims & sources

Scoped operating models

work context → merge-ready pull requestWork product review · Level 3

Summary and context

Summary

Homegrown one-shot coding agents that read work context, locate the right repo/workspace, implement a change end-to-end, and produce a PR for human review.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Over 1,300 completely minion-produced PRs merged per week in Part 2, with human review and no human-written code

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Stripe reports merged PR volume without an independent count or measurement window; Part 2 is explicitly later than Part 1.
Reported by
Stripe
Scope
Weekly merged PRs completely produced by Minions; Part 2 report, up from Part 1
Denominator
Unknown
Method
Unknown
Observation date
Unknown
Key observation

Over 1,000 completely minion-produced PRs merged per week in Part 1, with human review and no human-written code

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Stripe reports merged PR volume without an independent count or measurement window; no calendar metric date appears in the capture.
Reported by
Stripe
Scope
Weekly merged PRs completely produced by Minions; earlier Part 1 report
Denominator
Unknown
Method
Unknown
Observation date
Unknown

Architecture and primitives

Sandbox

Pre-warmed AWS EC2 devboxes in the QA environment, isolated from real user data, production services, and arbitrary network egress

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

Fork of Block's goose, orchestrated by code-defined blueprints that interleave agent loops with deterministic lint, git, and CI steps

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Curated subsets of Toolshed MCP tools for internal documentation, tickets, build status, and code intelligence; security controls constrain destructive actions

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Repository-scoped rule files shared with human-operated coding agents, plus internal context fetched through MCP

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

Full permissions inside quarantined devboxes; MCP security controls limit destructive actions, and production pull requests require human review

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Operating model evidence

Operating model assessment

Level 3 for work context → merge-ready pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source explicitly states that humans review and approve production pull requests.
Observation date
2026-02-20

Sources

  1. Minions: Stripe's one-shot end-to-end coding agents · Preserved Markdownhttps://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agentsengineering-blog · first-party · Last source verification: 2026-08-31
  2. Minions - Part Two · Preserved Markdownhttps://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2engineering-blog · first-party · Last source verification: 2026-08-31
  3. Hacker News discussion of Minions part two · Preserved Markdownhttps://news.ycombinator.com/item?id=47086557hn-thread · community · Last source verification: 2026-08-31
  4. Hacker News comment asking for examples and review evidence · Preserved Markdownhttps://news.ycombinator.com/item?id=47086907hn-comment · community · Last source verification: 2026-08-31
  5. Hacker News comment questioning the level of technical detail · Preserved Markdownhttps://news.ycombinator.com/item?id=47086953hn-comment · community · Last source verification: 2026-08-31
  6. Hacker News comment asking about shipped and maintained output · Preserved Markdownhttps://news.ycombinator.com/item?id=47087114hn-comment · community · Last source verification: 2026-08-31
  7. Hacker News comment questioning the value of pull request counts · Preserved Markdownhttps://news.ycombinator.com/item?id=47111453hn-comment · community · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
UberTask agent

Internal coding agent (unnamed)

Uber's internal coding agent, reported by its CTO as producing roughly 1,800 complete code changes per week.

Coding
coding request → complete code change: Unknown
Operating model, claims & sources

Scoped operating models

coding request → complete code changeUnknown · Level unknown

Summary and context

Summary

Uber's internal coding agent, reported by its CTO as producing roughly 1,800 complete code changes per week.

Fact · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Only limited public evidence supports this claim.

Reported metrics

Headline claim

~1,800 complete code changes per week (~8% of changes)

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Only limited public evidence supports this claim.
Reported by
Uber
Scope
Code changes written entirely by the internal coding agent and human reviewed
Denominator
All Uber code changes for the reported 8% share
Method
Unknown
Observation date
Unknown
Key observation

~1,800 complete code changes per week (~8% of changes at the time)

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Only limited public evidence supports this claim.
Reported by
Uber
Scope
Code changes written entirely by the internal coding agent and human reviewed
Denominator
All Uber code changes for the reported 8% share
Method
Unknown
Observation date
Unknown
Key observation

95% of engineers use AI tools monthly

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
Only limited public evidence supports this claim.
Reported by
Uber
Scope
Monthly use of AI tools among Uber engineers; not internal-agent-specific adoption
Denominator
Uber engineers
Method
Unknown
Observation date
Unknown

Architecture and primitives

Operating model evidence

Operating model assessment

Unclassified for coding request → complete code change; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The available source reports output volume but does not document where human attention returns.
Observation date
2026

Sources

  1. Uber's CTO on AI coding agents (Business Insider) · Preserved Markdownhttps://www.businessinsider.com/uber-cto-ai-coding-agentic-software-engineers-2026-3news · independent-secondary · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
UberBackground agent

uReview

An event-driven AI code reviewer for Uber's internal review platform that generates, grades, filters, deduplicates, and posts findings while leaving engineers in control of the reviewed change.

Code review
pull request → filtered AI review findings: Work product review
Operating model, claims & sources

Scoped operating models

pull request → filtered AI review findingsWork product review · Level 3

Summary and context

Summary

An event-driven AI code reviewer for Uber's internal review platform that generates, grades, filters, deduplicates, and posts findings while leaving engineers in control of the reviewed change.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Reported metrics

Headline claim

Uber's introduction reports reviews of over 90% of approximately 65,000 weekly diffs, with over 75% usefulness and over 65% addressed comments; a later paragraph says 65,000 diffs per month

Metric · Reported · Low confidence
Evidence and qualifications
Confidence reason
The opening paragraph reports approximately 65,000 weekly diffs, but the cost discussion says 65,000 per month. These conflicting periods remain unresolved; neither is independently verified.
Reported by
Uber
Scope
Introduction's reported weekly diff coverage, plus comment usefulness/addressed rates; later paragraph's monthly period conflicts and remains unresolved
Denominator
Introduction reports approximately 65,000 weekly diffs; cost paragraph says monthly. Usefulness covers rated comments; addressed rate covers posted comments
Method
Production coverage, developer ratings, and automatic addressed-comment detection
Observation date
2025-08
Key observation

Introduction reports reviews of over 90% of approximately 65,000 weekly diffs; cost discussion instead says 65,000 diffs per month

Metric · Reported · Low confidence
Evidence and qualifications
Confidence reason
The opening paragraph reports approximately 65,000 weekly diffs, but the cost discussion says 65,000 per month. These conflicting periods remain unresolved; neither is independently verified.
Reported by
Uber
Scope
Introduction's reported weekly diffs analyzed; monthly period in the cost paragraph conflicts and remains unresolved
Denominator
Approximately 65,000 weekly diffs according to the introduction; same volume described as monthly in the cost paragraph
Method
Unknown
Observation date
2025-08
Key observation

Over 75% of comments rated useful by engineers who interact with the tool

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Uber
Scope
Comments rated useful by engineers who provide feedback
Denominator
Comments with engineer interaction
Method
Useful / Not Useful rating links
Observation date
2025-08
Key observation

Over 65% of posted comments addressed in the same changeset

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Uber
Scope
Posted comments considered addressed in the same changeset
Denominator
Posted comments
Method
Five reruns on the final commit and semantic-similarity matching
Observation date
2025-08
Key observation

Median review latency of 4 minutes across all six Uber monorepos

Metric · Reported · Medium confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Reported by
Uber
Scope
Reviews across all six Uber monorepos
Denominator
Unknown
Method
Production latency telemetry
Observation date
2025-08
Key observation

Approximately 1,500 developer hours reportedly saved per week, based on an assumed 10-minute second review per processed commit

Metric · Reported · Low confidence
Evidence and qualifications
Confidence reason
The source estimates about 1,500 hours using over 10,000 commits and a 10-minute assumption; the rounded figures do not arithmetically reconcile exactly and are not observed time savings.
Reported by
Uber
Scope
Processed commits excluding configuration files; modeled second-review time savings
Denominator
Over 10,000 commits per week
Method
Processed commits multiplied by an assumed 10 minutes for a second human review
Observation date
2025-08

Architecture and primitives

Harness

A prompt-chained pipeline separates comment generation, confidence grading, validation, semantic deduplication, and category filtering; three specialized assistants were in operation when published

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

Reviews eligible code in Uber's six monorepos across Go, Java, Android, iOS, TypeScript, and Python; richer internal artifacts were not yet connected when published

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

Surrounding source context plus a shared registry of Uber-specific coding and style rules

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Context management

Comments and metadata are streamed through Kafka to Hive for feedback analysis, experiments, and operational dashboards

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Developer ratings, addressed-comment detection, and a curated benchmark tune prompts, thresholds, and models

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Combine prompts with deterministic filtering, deduplication, evaluation, and feedback instrumentation

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The article states this lesson directly and documents the pipeline.
Lesson

Roll out gradually by team and assistant while tracking precision, recall, usefulness, and false positives

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The article states this lesson directly and describes the rollout telemetry.

Operating model evidence

Operating model assessment

Level 3 for pull request → filtered AI review findings; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes AI-generated review findings while engineers remain in control of the change.
Observation date
2025-08-12

Sources

  1. uReview: Scalable, Trustworthy GenAI for Code Review at Uber · Preserved Markdownhttps://www.uber.com/us/en/blog/ureview/engineering-blog · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
WorkOSPlatform

Project Horizon

An internal autonomous 'code factory' where a continuously running swarm of agents handles the implementation loop while engineers focus on requirements and acceptance testing. Deliberately modular so the harness can evolve.

CodingCode reviewSecurity
requirements and acceptance criteria → tested implementation: Outcome review
Operating model, claims & sources

Scoped operating models

requirements and acceptance criteria → tested implementationOutcome review · Level 4

Summary and context

Summary

An internal autonomous 'code factory' where a continuously running swarm of agents handles the implementation loop while engineers focus on requirements and acceptance testing. Deliberately modular so the harness can evolve.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Sandbox

Cloudflare Containers + Sandbox SDK; disposable, tightly scoped sandboxes with explicit lifecycle APIs and egress controls; full monorepo stack in Docker dev containers

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Harness

Modular by design; the core article runs OpenCode in the sandbox; the Applied AI Showcase runs Claude Remote Routines. The harness is swappable as agent tech changes; separate PM, implementation, and prospective verification/security roles

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Tool access

A custom MCP server stitches internal data sources (Datadog, Sentry, Slack, WorkOS Pipes); all outbound traffic proxied through Workers with allowlists, limits, logging, and token injection

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Knowledge

AGENTS.md and CLAUDE.md capture scripts, docs, conventions; MCP codifies the patterns engineers already follow; Notion + Figma for specs/mockups

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Credentials

WorkOS Pipes (no OAuth/token-refresh to maintain); scoped short-lived GitHub tokens per user; engineers use their own identity in the MCP; least-privilege + egress controls

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

You need purpose-built agent infrastructure; a runtime you control end-to-end with lifecycle APIs and egress controls

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 4 for requirements and acceptance criteria → tested implementation; human attention boundary: outcome-review.

Inference · Catalog judgment · High confidence
Evidence and qualifications
Confidence reason
The source describes engineers focusing on requirements and acceptance testing while agents run the implementation loop.
Observation date
2026-05-06

Sources

  1. Project Horizon - an autonomous code factory at WorkOS · Preserved Markdownhttps://workos.com/blog/project-horizonengineering-blog · first-party · Last source verification: 2026-08-31
  2. Applied AI Showcase (Horizon with Claude Remote Routines) · Preserved Markdownhttps://workos.com/blog/applied-ai-showcaseengineering-blog · first-party · Last source verification: 2026-08-31
  3. An autonomous UI-quality program · Preserved Markdownhttps://workos.com/blog/autonomous-ui-quality-programengineering-blog · first-party · Last source verification: 2026-08-31
  4. Hacker News submission for Project Horizon · Preserved Markdownhttps://news.ycombinator.com/item?id=48039227hn-thread · community · Last source verification: 2026-08-31
Entry reviewed 2026-08-31
Y CombinatorPlatform

Internal agent infrastructure

Internal agent infrastructure and own harnesses built from the ground up, framed as making AI the operating system the whole organization runs on.

CodingOps
internal request → agent-assisted organizational work: Unknown
Operating model, claims & sources

Scoped operating models

internal request → agent-assisted organizational workUnknown · Level unknown

Summary and context

Architecture and primitives

Lessons and interpretation

Operating model evidence

Operating model assessment

Unclassified for internal request → agent-assisted organizational work; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence
Evidence and qualifications
Confidence reason
The source describes internal agent infrastructure but not a sufficiently specific human attention boundary.
Observation date
2026

Sources

  1. Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen) · Preserved Markdownhttps://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbookpodcast · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-09-09
ZupTask agent

CodeGen

A research-documented internal coding agent where constrained editing tools and layered safety controls mattered more than prompt tweaks.

Coding
constrained coding task → human-supervised edit: Continuous steering
Operating model, claims & sources

Scoped operating models

constrained coding task → human-supervised editContinuous steering · Level 2

Summary and context

Summary

A research-documented internal coding agent where constrained editing tools and layered safety controls mattered more than prompt tweaks.

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Architecture and primitives

Harness

Constrained editing tools; state-management concerns; progressive levels of human oversight

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.
Supporting component

Tools that limit what the agent can change, vs. unconstrained generation

Fact · Reported · High confidence
Evidence and qualifications
Confidence reason
A linked first-party source states the claim.

Lessons and interpretation

Lesson

Targeted tool design and layered safety controls matter more than prompt tweaks

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.
Lesson

Progressive levels of human oversight help build trust before granting autonomy

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The catalog derives this observation from the linked sources.

Operating model evidence

Operating model assessment

Level 2 for constrained coding task → human-supervised edit; human attention boundary: continuous-steering.

Inference · Catalog judgment · Medium confidence
Evidence and qualifications
Confidence reason
The research source describes progressive human oversight but does not establish background delegation.
Observation date
2026

Sources

  1. Building an Internal Coding Agent at Zup · Preserved Markdownhttps://arxiv.org/abs/2604.09805paper · first-party · Last source verification: 2026-08-31
Entry reviewed 2026-08-31