RoboHarness

A Simple Harness OutperformsVLA and World Action Models

Generalist robot manipulation with coding agents.

RoboHarness overview: a geometry tracker and minimal toolset connect a coding agent to a robot; an example rollout shows key insertion through visual feedback and geometric constraints.

Abstract

The rise of large language models (LLMs) has opened a path toward generalist embodied AI, raising the question of how LLMs can act in the physical world. One straightforward approach is training the model to generate actions: collect robot-action data and video demonstration, refine LLM architecture with an action head, and train a vision-language-action (VLA) model at scale.

However, we argue that coding agents have the potential to realize generalist embodied manipulation. All they need is a simple yet effective robot harness. We present RoboHarness: a robotic harness that gives LLM agents a visual-geometric control panel so that they can directly understand and invoke embodied tasks. With RoboHarness, Qwen3.8-Flash-Next outperforms state-of-the-art VLA models on BEHAVIOR Challenge 2025 tasks. The implementation is open source.

Method

Task results

Qwen3.8-Flash-Next · mean Q-score over five instances.

All task scores
Historical directory-reported Q-scores. The mean is over five instances per task.
Task301304306308310Mean Q
Turning on radio0.00000.00000.00001.00001.00000.4000
Picking up trash1.00000.33331.00001.00001.00000.8667
Putting away halloween decorations0.57140.71430.42860.71431.00000.6857
Cleaning up plates and food0.28570.28570.28570.14290.28570.2571
Setting mousetraps1.00001.00000.50001.00000.66670.8333
Hiding easter eggs1.00000.00000.77780.33330.00000.4222
Picking up toys0.00000.66670.50000.66670.00000.3667
Rearranging kitchen furniture0.25000.50000.25000.50000.50000.4000
Putting up christmas decorations inside0.33330.22220.33330.11110.11110.2222

BEHAVIOR Challenge 2025, with twice the challenge step budget. Task04 has no supplied archive. For task00 and task05, directory reports differ from some preserved evaluator JSON; the directory values shown here remain the reference.

Videos show the highest-scoring available recorded run for each task. Recorded-run Q-scores and the historical five-instance results are reported separately. Easter eggs has an interrupted recording with no final score. The recovered Halloween recording ends 6.2 seconds before its evaluation.

Evaluation details

Generalization across 100 objects

Objects

The catalog contains 100 rigid objects from the BEHAVIOR object library. Most have an extent of at most 0.21 m and a mass of at most 1.2 kg; the saucepan and one recorded outlier are retained as exceptions.

Task

The task is to pick up each object. Learned policies are evaluated in their training embodiment and simulation environment. For ASPIRE and RoboHarness, the object starts on the floor and the robot begins in a stooping pose. LIBERO-trained models use a cleared table with the object at its center.

Each trial runs for up to 2,000 simulation steps. Success means lifting the object at least 5 cm above its initial height.

The 100 rigid objects used to evaluate generalization, drawn from the BEHAVIOR object library.
100 objects from the BEHAVIOR library.

Generalization without robot data

Across the three tested LLMs, RoboHarness achieves 68–86 successes out of 100 without robot manipulation training data. More robot data does not consistently lead to better generalization among the VLA and world action model baselines.

Successes out of 100 objects versus estimated hours of robot data, comparing RoboHarness with VLA models, world action models, and a pick-function baseline.

Get started

RoboHarness is open source.

GitHub
Bash
git clone --recurse-submodules https://github.com/BinceQu/RoboHarness.git
cd RoboHarness
bash scripts/setup.sh