Embodied visual intelligence

RawVLA

Embodied Neural Image Signal Processor
for Robotic Manipulation

Shuhong Liu·Heng Zhou·Lingfeng Qian·Yuhao Fang·Xianbao Hou·Qianyu Zhou·Lin Gu·Wei Sui·Jianfei Yang·Ziteng Cui

UTokyo · D-Robotics · NTU · TohokuU · HKUSTGZ

An adaptive visual front-end that restores reliable perception for vision-language-action policies as real-world illumination shifts.

RAW ImagingBurst DenoisingWhite BalanceColor CorrectionTone CurveEmbodied ISPVision-Language-ActionIllumination RobustnessRawVLA-BenchReal-World Manipulation

01 / Overview

Robots should not lose
their skills when the light changes.

RawVLA places an embodied, adaptive neural ISP before a frozen VLA policy, translating RAW observations into policy-ready images without retraining the policy itself.

RawVLA overview and qualitative results across illumination
Figure 01 RawVLA adapts the imaging pipeline to support robust robotic manipulation across diverse illumination environments.

02 / ISP Perturbations

ISP choices substantially affect
VLA performance.

A controlled study on RAW images in simulation examines how conventional ISP components, including exposure control, denoising, white balance, color correction, tone mapping and quantization, affect task success across LIBERO and RoboTwin 2.0.

Six-axis ISP sensitivity results for exposure, tonal response, sensor noise, white point, color relation and bit depth on LIBERO and RoboTwin 2.0

LIBERO Six policy checkpoints expose large differences in tolerance, with exposure producing the sharpest collapse and world-action models showing narrower stable ranges.

RoboTwin 2.0 Qwen3-OFT and FastWAM remain sensitive at severe settings even after task-specific fine-tuning and visual domain randomization.

03 / Method

An embodied ISP that reasons
with the robot.

A streaming RAW burst is summarized into luminance and chroma descriptors. Recurrent states decode structured camera controls before the frozen VLA policy predicts actions.

Embodied imaging loop
168 FPS
RawVLA architecture with RAW burst encoder, recurrent states, structured ISP controls and frozen VLA policy

01Observe A causal linear RAW burst preserves information conventional pipelines discard.

02Adapt Factorized recurrent states track luminance and chroma over time.

03Embodied ISP Optimized by the downstream VLA action loss to generate images that favor accurate action prediction.

04 / RawVLA-Bench

Synthetic RAW sensing
across changing illumination.

RawVLA-Bench synthesizes sensor-space RAW observations across five illumination domains while preserving multi-view camera streams for controlled, reproducible evaluation.

5illumination domains
13,250evaluation rollouts
2,403paired training trajectories
408,075multi-view frames
RawVLA-Bench / 01

LIBERO

40 tasks · 10,000 paired evaluation rollouts

LIBERO RawVLA-Bench visualization across five illumination regimes, with Original, RAW and Default ISP observations for Agent and Wrist views
LIBERO · Benchmark Visualization Five matched illumination regimes show the original RGB reference, sensor-space RAW observation and fixed Default ISP output for Agent and Wrist views.
RawVLA-Bench / 02

RoboTwin 2.0

13 tasks · 3,250 paired evaluation rollouts

RoboTwin 2.0 RawVLA-Bench visualization across five illumination regimes, with Original, RAW and Default ISP observations for Agent and Wrist views
RoboTwin 2.0 · Benchmark Visualization The same five-regime protocol visualizes paired Original, RAW and Default ISP observations in a multi-camera dual-arm environment.

05 / Results

Results across simulation
and the physical world.

Compare every ISP baseline across five illumination regimes, then inspect task-level success rates from the real-world evaluation.

Interactive Result 05-A
Simulation success rates across illumination settings Interactive comparison of RawVLA and all ISP baselines on LIBERO and RoboTwin 2.0.

Result 05-B · Physical Robot

Real-World Evaluation

75.0% Normal Light 66.5% Low Light
Lighting / 01

Normal Light

TaskDefault ISPDarkISPRawVLA
Pick & Place966696
Close Drawer925294
Pick Flowers724466
Stack & Cover463844
Average76.5050.0075.00
Lighting / 02

Low Light

TaskDefault ISPDarkISPRawVLA
Pick & Place02292
Close Drawer01288
Pick Flowers01056
Stack & Cover0030
Average0.0011.0066.50

Success rate (%) · PI-0.5 · 50 trials per task

06 / Demos

From simulation
to the physical world.

Each benchmark uses independent rollout videos.

01

Simulation · Qwen3-OFT

LIBERO

Four independent method rollouts. Every player shows synchronized Agent and Wrist views from its rendered rollout.

Task: Pick up the butter and place it in the basket

65.18% overall success rate2× speed
Camera pipeline

Default ISP

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Neural ISP

RAM

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Neural ISP

RAWild

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Ours

RawVLA

Low Light
Agent ViewWrist View
Rendered Rollout✓ Success
02

Simulation · π₀.₅

RoboTwin 2.0

Four independent dual-arm rollouts, each preserving the Agent and Wrist views from the source videos.

Task: Lift the pot

54.98% overall success rate2× speed
Camera pipeline

Default ISP

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Neural ISP

DarkISP

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Neural ISP

RAW Adapter

Low Light
Agent ViewWrist View
Rendered Rollout× Fail
Ours

RawVLA

Low Light
Agent ViewWrist View
Rendered Rollout✓ Success
03

Physical robot · 1× speed

Real World

Every rollout combines a third-person iPhone view, a tabletop Agent view, and two Wrist views. Normal light is separated by task; low light is separated by method; disturbance recovery compares original and enhanced iPhone views.

66.50% low-light success rate4 tasks
Lighting / 01

Normal Light

RawVLA · Ours

Close Drawer

Normal Light
1× speed✓ Success
RawVLA · Ours

Pick & Place

Normal Light
1× speed✓ Success
RawVLA · Ours

Pick Flowers

Normal Light
1× speed✓ Success
RawVLA · Ours

Stack & Cover

Normal Light
1× speed✓ Success
Lighting / 02

Low Light

Camera pipeline

Default ISP

Low Light
Visualization Enhanced
1× speed× Fail
Neural ISP

DarkISP

Low Light
Visualization Enhanced
1× speed× Fail
Ours

RawVLA

Low Light
Visualization Enhanced
1× speed✓ Success
Lighting / 03

Low-Light Disturbance Recovery

Objects are repositioned during execution; RawVLA continues manipulation from the disturbed state.
Original shows the direct iPhone recording; Enhanced is post-processed only for demo visualization.

RawVLA · Recovery

Pick & Place

Original
1× speed✓ Recovered
RawVLA · Recovery

Pick & Place

Enhanced
1× speed✓ Recovered
RawVLA · Recovery

Pick Flowers

Original
1× speed✓ Recovered
RawVLA · Recovery

Pick Flowers

Enhanced
1× speed✓ Recovered
First page of the RawVLA paper

07 / Paper

RawVLA: Embodied Neural Image Signal Processor for Robotic Manipulation

Citation
@article{liu2026rawvla,
  title={RawVLA: Embodied Neural Image Signal Processor for Robotic Manipulation},
  author={Liu, Shuhong and Zhou, Heng and Qian, Lingfeng and Fang, Yuhao and Hou, Xianbao and Zhou, Qianyu and Gu, Lin and Sui, Wei and Yang, Jianfei and Cui, Ziteng},
  journal={arXiv preprint arXiv:2609.37530},
  year={2026}
}

RawVLA-Bench