Arena turns your raw data into agents, specialized at your task. Skip the reliance on someone else's frontier model.
Trusted by teams shipping agents in research, defense, finance, robotics and logistics.


Turn your workflows into training environments and let agents learn them directly.

Our evolutionary hyperparameter optimization and async-RL engine trains your agents 10x faster.
Reinforcement fine-tuning for LLM agents, with distributed training and one-click deployment.
Check your dataset or environment upfront and start training with full confidence.

Bring your data.
Use Arena to prepare it as LLM datasets and RL environments.
Validate everything with feedback before training begins.




Configure everything.
Select algorithms, rewards, constraints, and objectives.
Enable evolutionary tuning to explore promising configurations automatically.
Optimized async-RL engine and distributed training and across your entire compute stack.
Squeeze all the performance out of your spend.
Monitor metrics, sample efficiency, and checkpoints in real time.




One-click promote to production, securely hosted on your infrastructure.
Select checkpoints, track performance, benchmark against baselines.
Continual learning on live feedback.
Frontier models are generally good, but aren't specialized at what matters to you.
Agents stop learning after training and deployment.
Hosted fine-tuning and models carry a supply chain and data security risk.
Costly and time consuming training runs.
Tied to one frontier model provider where the cost scales with usage.
Small models fine-tuned on your task become experts.
Agents keep learning.
Live results feed the next run.
Training runs in your environment; model weights remain yours.
10x faster automated evolutionary training.
Use any open-source model, served on your own infrastructure.
Open-source framework with docs, examples, and community support
Single and multi-agent support across on/off-policy, offline RL, bandit and LLM training
Python-first, compatible with your data and environments
Works with your cloud compute and scales to multi-GPU
Training with open-source v2
Used by leading research
labs and institutions
Downloads from the community
Enabling breakthrough results in training AI agents for complex aerial interception missions with RTDynamics
Substantially cutting compute expenses and boosting training speed for RL workflows with Warburg AI
Dramatically increasing utilisation and reducing training time for complex bin-packing with Decision Lab
Bring a task, dataset, or environment. We'll train, tune, and deploy an agent on it with you in a live session.
Book a demo