We’re excited to announce Cortex Harness, a platform for running AI models on real robots with remote human guidance.
In our previous robotics evals, we tested how well frontier models could perform real-world manipulation tasks with robots. Those evaluations focused on individual attempts, leaving us with a broader question: What does it take for these models to keep doing useful work over time?
Current models are making progress on these tasks, but they still need help when they get stuck. Cortex Harness explores how timely human guidance helps robots recover and continue working. These interventions also provide recovery examples that could improve the models through further training, bringing us closer to our vision of agents moving from computers into the physical world.
We're also inviting researchers, developers, and industry partners to join our beta program. Participants will receive $100 in credits. Join the waitlist here.
Models act, humans intervene, recoveries become data
With Cortex Harness, operators can steer the robot with prompts or take direct control when the model needs help. The platform also collects these interventions as recovery data for future model training.
An operator starts by choosing a model and describing the task, then selecting the level of reasoning effort. The model picker supports large language models and vision-language-action models, including open-weight options. We use GPT-6 Astra with Ultrafast enabled for the tasks below.
As the robot works, the operator can watch its camera feeds and follow the model's reasoning and actions, making it easier to guide progress. This visibility helps the operator decide when a new prompt should guide the model's next attempt.
Steering by prompting
As the robot stacks cups, the operator prompts: "Use your free hand to get a better view." This lets the robot perform the task more accurately while the model stays in control.
Cortex Harness queues that instruction for the model's next planning step, so it can adjust the robot's camera view before trying again.
Watch steering by prompting in the full episode at 1× speed.
Remote Intervention: Operating a robot in San Francisco from Singapore
An operator in Singapore helps a robot in San Francisco pick up and place a battery into a slot, showing how remote assistance can support robotic work across distance and help robots complete tasks they cannot manage alone.
When the robot struggles to grasp or insert the battery, the operator remotely takes over with virtual reality controls, seamlessly helping the robot complete the task with minimal latency, then hands control back once the step is complete.
Using its current observations and the movements made during the intervention, the model decides what to do next and resumes the task.
Interventions become recovery data
Human intervention also creates an opportunity to collect useful data to improve models. Alongside autonomous rollouts, we can collect examples of where a model gets stuck, what a human corrects, and how the task continues:
Autonomous rollout → Failure/stuck state → Human correction → Recovery → Successful continuation
Instead of collecting only successful demonstrations from scratch, we can collect recovery examples from states that models actually reach during a task. This gives us more relevant training data, which we can curate into intervention trajectories for post-training and recovery training.
The goal is a continuous improvement loop: deploy, identify failures, recover with human help, collect recovery examples, train, and redeploy. This reduces the need for intervention over time and lets us measure that progress.
Watch remote intervention in the full episode at 1× speed.
What's next
This first version of Cortex Harness is just the beginning, and we will share more as we explore additional robots, tasks, and real-world environments.
If you’re interested in participating, join the Cortex Harness waitlist. Researchers, developers, and industry partners invited to early access will receive $100 in Cortex credits.
Join the waitlist here.