Blog

Why serious companies build their own agent platforms

Building at Turnkey
·
·
Conner Swann
Staff Engineer, Applied AI

About: Learn why Turnkey is building its own agent platform, what changes when agents move into shared company systems, and how identity, permissions, approvals, and auditability fit together.

Audience: Engineering leaders, AI platform teams, security teams, and developers building internal agents that need to work across company systems.

What you’ll learn:
  • Why companies move from developer agents to shared company agents
  • Why identity and authentication become harder once agents leave the developer’s laptop
  • How permissions need to map across systems like GitHub, Linear, Slack, and AWS
  • Why background agents need durable human approvals
  • How to preserve attribution so every agent action has a clear owner
  • Why shared agents need a common layer for credentials, policy, approvals, and audit logs

Reading time: ~8 minutes

In January, Ramp published Why We Built Our Own Background Agent, an architecture post about its internal coding agent, Inspect.

At publication, Inspect wrote roughly 30 percent of the pull requests merged into Ramp's frontend and backend repositories. By April, Ramp said Inspect was writing about 70 percent of all merged pull requests, with use extending beyond its engineers. A June Inspect at Scale retrospective reported that Inspect created 76.5 percent of all merged pull requests over the prior six months.

Ramp did not require its engineers to use it. Inspect earned that share by giving each task a remote development environment, access to internal systems, and enough tooling to verify backend and frontend changes before it returned the work.

Inspect's growth from 30 percent to more than 70 percent of merged pull requests shows that it became part of Ramp's normal development process. Ramp chose to build around its own workflows and systems even though equipping every engineer with Cursor, Claude Code, or Codex would have initially cost less. Stripe made a similar investment with Minions, as did Block with Goose, Shopify with River, and now Turnkey with Valet.

Agent infrastructure

Why build your own agent infrastructure?

Once an agent acts in shared company systems, the hard part is no longer the model.

Off-the-shelf agent
  • Great on one laptop
  • Inherits local access
  • Waits for its developer
Shared systems introduce hard questions
  • Who requested the action?
  • What are they allowed to do?
  • Can approval survive a restart?
  • Who owns the final change?
Company agent platform
  • Narrow authority
  • Durable approvals
  • Auditable attribution
Common layerIdentityPolicyCredentialsApprovalsAudit logs

Building Valet, Turnkey's internal AI agent platform

At Turnkey, we built Valet to handle a similar class of work as Ramp’s Inspect.

Valet is our internal agent platform for tasks that span code and company systems. Each user has a personal Valet that remembers context, runs scheduled work, and starts isolated sessions with full development environments. A session can receive requests from the web, Slack, or Telegram.

Valet uses existing models, such as Claude’s Opus, and gives them a controlled way to act. A session can clone a repository, run tests, open a pull request, file a Linear issue, or use another company integration under the requester's identity. The platform keeps credentials out of the model's context, pauses risky actions for human review, and records who requested each action.

For companies that want to use these models, having a model just running in one isolated machine is problematic. Giving it authority inside a company creates identity, authorization, policy, and audit challenges. Those problems are specific to the company, connected to every system it uses, and difficult to buy in as a product from an outside vendor.

How companies build internal AI agents, in three tiers

Company agents

The three tiers of company agents

Personal productivity → shared automation → accountable infrastructure.

Tier 1:Developer's agentOn the laptop
  • Easy to buy and adopt
  • Inherits local environment
  • Great for one developer
Not shared company infrastructure
Tier 2:Shared botOne identity for everyone
  • One shared entry point
  • Broad combined access
  • Bot becomes the author
Scales to headcount, then incidents without owners
Tier 3:Company agentAccountable infrastructure
  • Narrow, scoped authority
  • Human or automated principal
  • Policy + approvals + audit
Shared agent, without superuser power

The platform becomes necessary when an agent can act in systems other people depend on.

Most companies doing serious work with agents move through three tiers, usually in order.

Tier 1: The developer's agent

The first tier is the agent on a developer's laptop. Cursor, Claude Code, Codex, and similar tools live where the developer works. They inherit the developer's environment: browser sessions, SSO cookies, gh auth login, AWS credentials, local configuration, and whatever else is already present.

These tools are easy to buy, fast to adopt, and useful. They work well for the person who owns the laptop. They do not try to become a shared company infrastructure.

Tier 2: The shared bot

The second tier is the shared bot. A team connects one Slack app or scheduled agent to a service account and gives it a set of OAuth tokens. Everyone can now ask the same agent to do work. This setup is quick to prototype and gives the company one shared entry point.

In this setup, the bot becomes the author of every action, regardless of who requested it. Its credentials usually cover the combined access needed by every requester. It may be able to read an executive's messages while answering an intern's question.

Audit records identify the bot instead of the requester. Revoking the requester's account does not remove any of the bot's permissions. Tier two scales to headcount for about six weeks, then starts producing incidents that do not have owners.

Tier 3: The company agent

The third tier replaces bot-owned authority with authority from an authenticated human or documented automated principal. The shared agent runs on company infrastructure, uses company-managed credentials, applies company policy, and leaves records the company can audit. This requires a credential store, a policy engine, and an audit log.

For internal AI agents, auth is the whole problem

An IDE agent inherits identity from the developer's laptop. This is what makes these agents so attractive and easy to set up. A company agent runs without that authenticated environment. Its platform must provision, scope, rotate, resolve, and revoke every credential the agent uses.

Consider my personal Valet. It may need GitHub access limited to repositories I can reach, Linear access for teams where I can create issues, and Slack access to my own messages.

A scheduled session may need some of that authority at 3 a.m., when no human is present to complete an approval flow. Putting long-lived tokens in environment variables is easy, but it also leaks authority into child processes, crash reports, and any tool that can read the environment.

Valet resolves credentials when a tool runs. The model receives a tool, not the underlying secret. The platform can mint short-lived access from a central credential store, attach the authenticated user, and revoke access without providing access to where the credentials are at rest.

Ramp reached a similar conclusion when it designed its agent identity model for agents that manage money. Ramp ties each agent to a human sponsor, limits the agent to a subset of the sponsor's permissions, records both actors, and gives the access an explicit expiration and revocation lifecycle.

Turnkey works on the same identity problem for customers: how humans and machines receive narrowly scoped authority without exposing the underlying keys. Valet applies that approach to our own agents.

Auth flow

How Valet auth works

Identity follows the request. Secrets never enter the model context.

Person

Requests work in Slack, web, or Telegram

Valet session

Preserves requester and task context

Policy resolver

Checks requester, action, arguments, and context

Credential store

Mints short-lived, scoped access

Tool

GitHub, Linear, Slack, AWS

The model gets a tool interface, not the underlying secret.

Valet recordsRequesterPolicy decisionApprovalCredential useResult

AI agent permissions do not map cleanly across systems

After Valet identifies the requester, it still has to determine what that person is allowed to do in each connected system.

Every system handles permissions differently. GitHub has organizations, repositories, teams, installations, and repository roles. Linear has workspaces, teams, and members. Slack has workspaces, channels, private conversations, and enterprise organizations. AWS has accounts, organizational units, policies, and roles.

A useful company agent has to work across all of them.

Suppose I tell Valet in Slack, "File a bug in the repository we were just looking at."

Valet has to connect the Slack conversation to my identity, figure out which GitHub organization and repository I mean, and verify that I have permission to act there. It then has to map that repository to the Linear team that owns it and confirm that I can create an issue for that team.

Those checks also need to happen when the tool actually runs, not just when the model suggests the action.

No connected service has the full permission picture on its own. A company can build this common policy layer itself or use a unified identity platform, but either way it needs to translate each service's permissions into a shared model and keep those mappings up to date.

At Turnkey, we make the authorization decision when an agent invokes a tool. The policy evaluates the requester, the action, its arguments, and the execution context together. We use this approach for our customers' agents, and Valet applies it to our own.

Background agents need durable human approvals

A background agent is useful because it can keep working after the requester leaves. A task might start from Slack, a schedule, or an event and complete many routine steps on its own.

But if the agent later reaches an action that needs human approval, it should be able to pause there without requiring someone to supervise the entire task.

At each tool call, Valet evaluates who requested the work, what the agent wants to do, the arguments it supplied, and the current context. Policy can allow or deny the action immediately. It can also require approval from the requester.

When approval is needed, Valet stores the proposed action and pauses the tool call. The requester can review it from an available channel, then approve or deny it.

Because the pending decision is stored outside the running process, it survives a restart. Valet resumes the tool call only after receiving an answer.

Every integration and interface uses the same policy resolver. A request may start in Slack and continue in a sandbox, but the action receives the same policy decision wherever it runs.

Every agent action needs an owner and an audit trail

When an agent changes a company system, the record should show whose authority it used.

Naming the agent is not enough. The agent performed the action, but a person or automated process requested the work and gave it permission to act.

Valet carries that identity from the original request through policy evaluation, credential resolution, and tool execution. It records what the agent proposed, how policy evaluated it, whether a person approved it, and what the tool returned.

If the work produces a git commit, for example, the sandbox authors the commit as the requester.

This history matters because a single task can move across several systems over time. It might begin with a scheduled prompt, continue in a sandbox, and finish in GitHub after the requester has left.

Valet preserves the link between the original requester and the final change, even when no person is present for the last step.

Final thoughts: Shared agents need shared infrastructure

These constraints are not specific to Turnkey, Ramp, or software development. They appear when an agent moves from one person's machine into shared company systems.

If an agent only edits local files and waits for its developer, buy the best tool for that job. If a shared bot will answer low-risk questions for a short experiment, build the bot and learn from it. Both choices are reasonable.

Once an agent can open a pull request, file a ticket, page an on-call engineer, read private messages, or move money, a sandbox is insufficient. The company will also need per-person authority, a common policy boundary, durable approval requests, and end-to-end attribution. That’s tier three.

Companies keep building agent platforms because off-the-shelf coding agents do not provide this. A shared operating layer for permissions and authentication lets an agent use company systems without becoming an unaccountable superuser.

In the next post, I will cover six months of running Valet v0 in production: which product bets held up, which runtime choices did not, and why we rewrote the substrate from scratch. The final post will cover Valet v1 and how to run it yourself.

Conner Swann
Staff Engineer, Applied AI

Related articles

Why we built Swaps and Earn: Onchain revenue without the engineering burden

Learn why Turnkey built Swaps and Earn, how each product works, and how applications can turn onchain trading and yield into revenue.

Here’s how (and why) we built native MFA and Scoped Sessions at Turnkey

Why authentication gets harder to enforce at scale and how Turnkey’s MFA and Scoped Sessions bring these controls directly into the API.

Building at Turnkey
September 9, 2026