BACK
10/5/2026

Introducing AgentEnv: An Open-Source Framework for Building RL Environments

By Edgar Arakelyan, Tejas Polakam, Yzabelle Go

An agent is only as good as the world it operates in, and building worlds that are realistic, accurate, and scalable all at once is the hard part of scaling RL. Every RL environment Scale builds runs on AgentEnv. Today we're open-sourcing it.

Explore AgentEnv · View on GitHub

RL environments are getting harder to build

Scaling RL takes a steady supply of tasks for models to practice on, and those tasks live inside RL environments: simulated slices of the real world with the tools, data, and rubrics/verifiers an agent needs to attempt a task and learn from it.

Unlike pre-training, which can learn from large corpora of public and private data, RL has no equivalent corpus. RL environments and their tasks are mostly built from scratch, and the cost of building them grows linearly with their fidelity.

Designing and building RL environments takes close collaboration across different people: engineers who build the environment, ML researchers who design the tasks and rewards, and domain experts who know what good work looks like. AgentEnv is built with this modularity in mind so each of them can own their piece in parallel.

The result is a three-way constraint to what makes a great RL Environment. Environments have to be realistic enough that behavior transfers into valid learning signals, accurate such that the reward reflects the agent’s true behavior rather than noise in the setup, and scalable enough to support different tasks/goals and trajectories.

Building Blocks and Interoperability

Most tools for agent training start from the task and build a world around it. AgentEnv’s philosophy starts from the world first, so one Environment can host many Tasks and many agents. AgentEnv is built on two core principles

Building blocks: What an Environment is made of. An Environment is built from versioned composable blocks: Environments like Slack or Gmail, an Artifact of data that populates them, and Rules (the clock, triggers, and roles) that govern how it behaves.

Interoperability: What an Environment works with. An Environment doesn't assume anything about who or what runs it. Any agent can connect to it, any sandbox can host it, and your cloud can run it, all through open protocols instead of custom glue.

Dynamic, Composable Environments

Environments are building blocks: In AgentEnv, an Environment is anything an agent can act on: Slack, Gmail, a help desk, a browser, a desktop program like Blender, or even a video game. Each Environment contains an Environment Card as part of the open AgentEnv Environment Protocol; a manifest of what the Environment is, which tools it offers and how to reach them.

Figure 1. (Illustrative) The Environment Card: on the left as you read it, on the right as agents read it at /.well-known/agent-env.json. An agent reads the card, then calls email_search over MCP.

Building blocks snap together: Combine a help desk, email, and customer records environments, and you get a complete Support Desk Environment. An accounting firm, an enterprise company, and a SupportDesk environment can all share the same underlying Slack and Email environments, but each will be loaded with its own data Artifact, alongside any additional environments they need. At Scale, we build entire simulated workplaces this way, with a dozen or more environments working together.

Figure 2. Shared environments are built and tested once. A fix or improvement to the simulated Slack reaches every environment that uses it.

Designed for Interoperability

Agent Agnostic. An Environment's card tells an agent which tools it serves and how to reach them, over MCP, REST or a generated CLI. An agent's Agent Card, from the open A2A protocol, does the same in the other direction.

Pair any Environment Card with any Agent Card and the agent can start work, with no custom wiring. The same Task runs unchanged and consistent on Claude Code, OpenAI's agents, or any of your own custom agents.

Figure 3. The same Environment and Task run unchanged against any agent, in any sandbox. Four sandbox providers ship with the framework, and you can register your own by name.

Infra Agnostic. Build an Environment on your laptop and it runs unchanged weather in your cloud or local. AgentEnv comes with native, day-one support for AWS [S3, ECR, Secrets Manager], Google Cloud [Cloud Storage, Artifact Registry, Secret Manager] all set through a single config.toml file. In the future, we will be adding more providers for Sandboxes and Cloud support

Sandboxes are just as pluggable: AgentEnv supports four OOTB sandboxes (local Docker, Modal containers, and Modal and E2B virtual machines), and you can register your own as well.

At Scale, Modal is our default sandbox provider supporting all environments, task runs, evals and more!

Build any AgentEnv Task, step by step

Now that we can define any world as an AgentEnv Environment, the next question is what you can do in it. The answer turns out to be almost anything. Since any slice of the real world can be represented as an RL environment, almost any goal, task, or workflow can be represented as an AgentEnv Task.

A Task is constructed as a DAG of AgentEnv steps. Each step represents a primitive action on an Agent, the Environment and more, like setting up the world, loading data, setting the clock, bringing in an agent, handing it the work, collecting what it made, grading the result with a rubric or a strict check.

Because steps snap together the same way Environments do, an AgentEnv Task isn't limited to being a traditional RL training sample.

One Task might give an agent a written brief and a sandbox with Blender, and ask it to build and render a 3D scene. There are no rubrics or verifiers; the Task is simply a way to experiment without managing the application infrastructure yourself.

Another Task might drop an agent into our support desk environment to resolve a refund for a customer named Dana, then grade its trajectory and the outcome.

Every AgentEnv step you build can be reused across any task!

Figure 4. A customer refund and a 3D modeling job look nothing alike, but both are AgentEnv Tasks built from the same kinds of steps.

What comes out of the box

AgentEnv ships 48 built-in step types, covering everything from deploying an Environment to setting its Rules of the World to grading trajectories. Most RL Tasks can be built from these alone, with no new code. For the rest, you can add your own steps as plugins.

Example plugins

Plugins allow you to extend and build on top of AgentEnv and share with the community.

OpenCiv3 RL Env

AI agents play the open-source Civilization III remake against each other, messaging in public as they play.

Minecraft/Catan/Poker RL Env

Frontier models play poker, CATAN and Minecraft with and against each other.

iOS Mobile RL Env

AI agents drive a real iPhone in the cloud, reading the screen, tapping, swiping and typing through MCP tools.

Get started today

We want to hear all about what you build! File bugs and ideas as GitHub issues, and if you make an environment, sandbox, or step others could use, package it as a plugin and share it.

Explore AgentEnv · View on GitHub