Coarena by Coasty takes a practical approach to evaluating AI agents that can actually use a computer. Instead of relying on a fixed collection of benchmark questions, it puts two computer-use agents against the same real-world task and lets people judge the results without knowing which model produced which answer.
The idea is particularly useful for anyone trying to understand how AI performs beyond text generation. A computer-use agent has to navigate websites, interact with interfaces, follow instructions, recover from unexpected situations, and complete a task from beginning to end. Watching those attempts side by side makes the differences much easier to understand.
The platform describes itself as an arena for real-world evaluations, combining live tasks, blind human judgment, published metrics, and a dataset designed to make the evaluation process inspectable rather than mysterious.
The interface is intentionally centered around the battles themselves. Visitors can move between the arena, leaderboard, benchmark, dataset, and supporting evidence without having to navigate through an overloaded dashboard.
The battle format is especially easy to understand. Instead of reading a technical report about an agent, a visitor can watch two attempts and form an opinion about which one actually handled the task better. That makes the product approachable even for people who are interested in AI evaluation but are not machine-learning specialists.
The strongest part of the evaluation model is its emphasis on evidence. Rankings are not presented as isolated numbers. The methodology explains how ratings are calculated, how confidence intervals are handled, which battles qualify for ranking, and why certain results may be excluded.
The system uses a Bradley-Terry statistical model for its published ranking and reports confidence intervals alongside ratings. This is a useful distinction because a model with only a handful of battles should not be treated as equally reliable as one with a much larger evaluation record.
Another thoughtful detail is the use of blind judging. If a reviewer knows which model is responsible for an answer, expectations can easily influence the result. Hiding model identities during evaluation creates a cleaner comparison between the actual outcomes.
The platform is built around computer-use agents rather than conventional chatbots. Its evaluations can therefore reveal abilities that are difficult to capture with ordinary text benchmarks.
Agents can be evaluated on practical browser and desktop interactions, with the same action space and task conditions applied to both sides of a battle. The platform also records the trajectories of runs, allowing evaluation data to contain more than a simple win-or-loss label.
For researchers and developers, the dataset and metrics API add another useful layer. Instead of treating the leaderboard as the final destination, technically minded users can investigate the evidence behind the published measurements.
Privacy deserves particular attention because computer-use agents interact with visual computer environments. The platform states that screenshots are not included in its licensed datasets, while stored page captures shown in replays are subject to masking controls.
At the same time, users should understand an important limitation before submitting sensitive information. During an active run, the model provider operating the agent receives the raw, unmasked screen captures required for the agent to see and interact with the computer. Users should therefore avoid directing tasks toward private accounts, passwords, confidential documents, or other sensitive information.
Attached files are also sent to the agents running the task and retained with the task. This makes the platform transparent about an important part of its architecture, but it also means users should think carefully about what they upload.
One of the most interesting applications is AI research. Developers can use the arena to compare computer-use systems under practical conditions rather than relying entirely on synthetic benchmarks.
It can also be useful for AI enthusiasts who want to see how different frontier agents behave when faced with the same challenge. A replay can reveal small but important differences: one agent may find a shortcut, another may recover better after an error, while a third may produce a more reliable final result.
AI companies and researchers can potentially use the published dataset and metrics to study agent behavior, evaluate new models, and understand where computer-use systems still struggle.
For students and people learning about AI agents, the battle format offers an unusually visual way to understand evaluation. Instead of starting with statistical terminology, they can watch the agents work and then explore the methodology behind the scores.
No conventional paid subscription plans or public pricing tiers are presented as the main offering on the current website. Watching battles is open, while posting tasks and judging battles require an account.
The platform is currently structured more like a public evaluation arena than a conventional SaaS product with Free, Pro, and Enterprise packages. Users interested in participating can create an account and explore the available battle experience, while researchers can also access the published evaluation resources.
Start by opening the arena and exploring the available battles. Watching does not require participation, so visitors can immediately see how computer-use agents approach real tasks.
To participate, sign in with a supported account. From there, users can submit tasks for agents to attempt or judge existing battles. During judging, the identities of the competing agents are withheld so the decision can focus on the quality of their work.
For a deeper analysis, explore the leaderboard and benchmark sections after watching several battles. Researchers can go further by examining the available dataset and metrics resources rather than relying only on the headline ranking.
A good first experience is to watch several battles involving different types of tasks. After a few comparisons, the differences in navigation, error recovery, instruction following, and final results become much easier to recognize.
Traditional AI benchmarks usually provide a predetermined set of questions or tasks and measure how models respond. That approach works well for many language and reasoning problems, but it can become less representative when evaluating agents that operate websites and computer interfaces.
This platform takes a different route by continuously evaluating agents through an arena model. The task is the same for both competitors, the identities are hidden during judgment, and the resulting battles contribute to an evolving evaluation record.
Another difference is transparency. Rather than presenting a leaderboard without explaining how its numbers were produced, the platform publishes detailed rules covering matchmaking, rating calculations, confidence intervals, judging, and exclusions. That makes it particularly interesting for people who care about the methodology behind an AI ranking, not just the ranking itself.
Computer-use AI is moving beyond answering questions and toward actually operating software, websites, and digital environments. That shift creates a need for evaluation methods that measure what agents can accomplish in realistic situations.
This platform offers a compelling answer by turning those evaluations into head-to-head battles that people can watch and judge. The combination of real tasks, blind comparison, statistical ratings, replays, published methodology, and accessible evaluation data gives it a distinctive place in the growing AI benchmark landscape.
For anyone following the progress of computer-use agents, it is worth spending some time with the battles themselves. The numbers are useful, but seeing two agents attempt the same task often tells the more interesting story.
It is designed to evaluate AI agents that can use computers, browsers, and digital interfaces to complete real-world tasks.
No. The evaluation system is designed to keep the identities of competing agents hidden while a judgment is being made, helping reviewers focus on the actual results.
Yes. Watching the available battles is open to visitors. An account is required for activities such as posting tasks and judging battles.
The published ranking uses a Bradley-Terry statistical model and includes confidence intervals. The site also distinguishes its statistical ranking from its live Elo-style ticker.
Yes. A dataset and metrics API are provided as part of the evidence section, giving researchers additional ways to examine the evaluation results.
No. Users should avoid uploading confidential or sensitive information. Files attached to tasks are sent to the agents running those tasks, and the model providers involved in a computer-use run receive the screen information required to operate the agent.
The current website does not present conventional subscription pricing tiers as its primary offering. Watching is open, while participation requires an account.
The service is built, owned, and operated by Coasty Systems, Inc., and the company states that the product is backed by Y Combinator.
Instead of relying on a fixed test set alone, the arena evaluates agents through real tasks, head-to-head runs, blind human judgment, replays, and a continuously developing evaluation record.
AI Testing & QA , AI Research Tool , Large Language Models (LLMs) , AI Developer Tools .
These classifications represent its core capabilities and areas of application. For related tools, explore the linked categories above.