Simulation · Qwen3-OFT
LIBERO
Four independent method rollouts. Every player shows synchronized Agent and Wrist views from its rendered rollout.
Task: Pick up the butter and place it in the basket
Embodied visual intelligence
An adaptive visual front-end that restores reliable perception for vision-language-action policies as real-world illumination shifts.
01 / Overview
RawVLA places an embodied, adaptive neural ISP before a frozen VLA policy, translating RAW observations into policy-ready images without retraining the policy itself.
02 / ISP Perturbations
A controlled study on RAW images in simulation examines how conventional ISP components, including exposure control, denoising, white balance, color correction, tone mapping and quantization, affect task success across LIBERO and RoboTwin 2.0.
LIBERO Six policy checkpoints expose large differences in tolerance, with exposure producing the sharpest collapse and world-action models showing narrower stable ranges.
RoboTwin 2.0 Qwen3-OFT and FastWAM remain sensitive at severe settings even after task-specific fine-tuning and visual domain randomization.
03 / Method
A streaming RAW burst is summarized into luminance and chroma descriptors. Recurrent states decode structured camera controls before the frozen VLA policy predicts actions.
01Observe A causal linear RAW burst preserves information conventional pipelines discard.
02Adapt Factorized recurrent states track luminance and chroma over time.
03Embodied ISP Optimized by the downstream VLA action loss to generate images that favor accurate action prediction.
04 / RawVLA-Bench
RawVLA-Bench synthesizes sensor-space RAW observations across five illumination domains while preserving multi-view camera streams for controlled, reproducible evaluation.
40 tasks · 10,000 paired evaluation rollouts
13 tasks · 3,250 paired evaluation rollouts
05 / Results
Compare every ISP baseline across five illumination regimes, then inspect task-level success rates from the real-world evaluation.
Result 05-B · Physical Robot
| Task | Default ISP | DarkISP | RawVLA |
|---|---|---|---|
| Pick & Place | 96 | 66 | 96 |
| Close Drawer | 92 | 52 | 94 |
| Pick Flowers | 72 | 44 | 66 |
| Stack & Cover | 46 | 38 | 44 |
| Average | 76.50 | 50.00 | 75.00 |
| Task | Default ISP | DarkISP | RawVLA |
|---|---|---|---|
| Pick & Place | 0 | 22 | 92 |
| Close Drawer | 0 | 12 | 88 |
| Pick Flowers | 0 | 10 | 56 |
| Stack & Cover | 0 | 0 | 30 |
| Average | 0.00 | 11.00 | 66.50 |
Success rate (%) · PI-0.5 · 50 trials per task
06 / Demos
Each benchmark uses independent rollout videos.
Simulation · Qwen3-OFT
Four independent method rollouts. Every player shows synchronized Agent and Wrist views from its rendered rollout.
Task: Pick up the butter and place it in the basket
Simulation · π₀.₅
Four independent dual-arm rollouts, each preserving the Agent and Wrist views from the source videos.
Task: Lift the pot
Physical robot · 1× speed
Every rollout combines a third-person iPhone view, a tabletop Agent view, and two Wrist views. Normal light is separated by task; low light is separated by method; disturbance recovery compares original and enhanced iPhone views.
Objects are repositioned during execution; RawVLA continues manipulation from the disturbed state.
Original shows the direct iPhone recording; Enhanced is post-processed only for demo visualization.
07 / Paper
@article{liu2026rawvla,
title={RawVLA: Embodied Neural Image Signal Processor for Robotic Manipulation},
author={Liu, Shuhong and Zhou, Heng and Qian, Lingfeng and Fang, Yuhao and Hou, Xianbao and Zhou, Qianyu and Gu, Lin and Sui, Wei and Yang, Jianfei and Cui, Ziteng},
journal={arXiv preprint arXiv:2609.37530},
year={2026}
}