AeroManip-VLA

Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations

Rui Huang1 Yanlin Mu2,1 Lidong Li1 Yucong Wang1 Zichen Yan1 Lin Zhao1

NUS logo1National University of Singapore BIT logo2Beijing Institute of Technology

80K+
RL-generated demonstrations
no human teleoperation
5
parameterized skills
Pick · Place · Nav
Open · Close
3,900+
parallel environments
on a single RTX 5090
6.7×
throughput vs. AIR-VLA
at comparable GPU memory
4.2×
faster picks with RL
3.8 s vs. 15.8 s (expert)

01 · Video

Full Demo Video

All demonstrations on this page, compiled into one video.

02 · Overview

Aerial VLA data at scale, without teleoperation

AeroManip-VLA is a GPU-parallel benchmark for aerial manipulation. Reusable RL policies and expert task rules generate demonstrations automatically, while every rollout is annotated for task progress and flight safety.

2×

Aerial pick-and-place across diverse apartment scenes and objects.

Overview of the AeroManip-VLA framework
AeroManip-VLA at a glance. (a) 80K+ RL-generated expert trajectories across apartment scenes and six task types. (b) A VLA backbone consumes three onboard views (FPV + two gripper cameras), a language instruction and proprioception, and outputs action chunks of [vx, vy, vz, yaw, gripper]. (c) Long-horizon aerial manipulation decomposed into approach, align, pick, navigate and place.

GPU-parallel aerial VLA benchmark

Massively parallel simulation with low-level, payload-aware flight and manipulation control, built on ManiSkill3 and ReplicaCAD.

Demonstrations without teleoperation

Reusable RL policies combined with expert task rules generate data across randomized objects, scenes and initial poses, from basic skills to long-horizon tasks.

Structured analysis & baselines

Automated event labeling and trajectory categorization expose task progress and safety failures; ACT, Diffusion Policy, π0 and π0.5 are benchmarked.

03 · Benchmark

Benchmark Design

Four task settings built from five parameterized skills — PickPlaceOpenCloseNav — executed continuously within a single episode, with grasping realized through simulated physical contact (no attachment or teleportation).

TidyHouse

Object-specific Pick and Place for household rearrangement.

PrepareGroceries

Pick and Place of grocery items for kitchen rearrangement.

PackageDelivery

Outdoor: pick a package from a vehicle, fly to the door, place it on a cabinet.

LongHorizonRearrangement

Long-range navigation, pick-and-place, and opening / closing cabinets and fridges.

Two aerial manipulation platforms
Two X500-based platforms: a lightweight 1-DoF downward gripper (left) and a forward-reaching 3-DoF arm (right), both with UMI-style fingertips. Both share a payload-aware low-level controller and PX4-style control allocation, keeping the simulated flight stack close to real hardware for easy sim-to-real transfer.
Observation configuration
Observations: one FPV and two gripper-mounted RGB-D cameras (128×128) plus a 19-D proprioceptive state. The third-person view is for visualization only.

04 · Demonstrations

Expert Demonstrations

A privileged hybrid expert combines frozen learned components, rule-based guards, geometric planning and scripted feedback to produce reliable demonstrations in parallel environments. Each tile shows the third-person view (top), onboard RGB (middle) and depth (bottom).

Pick — descend onto the target object and grasp it through physical contact.

RL: Pick (TidyHouse)

RL: Place (TidyHouse)

05 · Reinforcement Learning

Extensive RL

Skill policies are further optimized with PPO in GPU-parallel environments. RL costs far more interaction than behavior cloning, but discovers more diverse and much faster successful behaviors.

67–81k vs. millions

Demonstration transitions BC needs to approach expert Pick success, vs. environment interactions PPO needs.

283 vs. 46

Occupied spatial cells of successful PPO approach paths vs. the expert on PrepareGroceries (205 vs. 53 on TidyHouse).

3.8 s vs. 15.8 s

Mean time to a successful pick for PPO vs. the expert and BC on TidyHouse, mostly from a shorter approach.

Learning efficiency of BC vs PPO
BC is substantially more interaction-efficient than PPO, but relies on expert demonstrations. BC markers report held-out success after 1, 2, 4 and 8 collection rounds; PPO curves show training success from scratch; dotted lines denote the expert.
Trajectory diversity
PPO produces substantially more diverse successful approach trajectories. Expert and BC follow a near-deterministic L-shaped approach, whereas PPO spreads across the X–Z plane, with higher path deviation, space coverage and velocity-direction entropy.
Execution time
PPO discovers diverse yet substantially faster trajectories. Mean time to success decomposed into approach, align, grasp, lift and hold phases (left) and the distribution of successful episode lengths (right).

06 · 3-DoF Arm

Articulated & Long-Horizon Tasks

With a forward-reaching 3-DoF arm, the aerial manipulator opens and closes fridges and kitchen counters, and chains these interactions with pick, navigation and place in a single continuous episode.

Open & Close

Open fridge

Open kitchen counter

Close fridge

Close kitchen counter

Open → Pick → Navigate → Place

“Open the drawer, take out the bowl, and place it on the dark brown table.”

“Open the fridge, take out the apple, and place it on the dark brown table in the living room.”

07 · Outdoor

Package Delivery

1Pick the package from the van roof
2Fly to the designated house
3Place it on the porch table
4×

08 · Annotation

Event Annotation & Failure Analysis

Every rollout is labeled with task-specific interaction events and aerial safety constraints — excessive tilt, altitude violations, non-target contact, payload loss, platform stability and clearance — so success is judged jointly with the physics of flight, and failures are categorized rather than just counted.

1.5×

Success The released object comes to rest inside the goal region (green) with only mild contact.

1.5×

Failure Cumulative contact force exceeds the limit after release — flagged as a release collision.

Failure cases

Pick · tuna fish can

Place · bowl

Long pick-and-place · tuna fish can

09 · Results

Policy Benchmarking

We evaluate visuomotor imitation learning (ACT, Diffusion Policy) and pretrained VLAs (π0, π0.5) on PrepareGroceries. Pretrained VLAs transfer best, yet Place is consistently easier than Pick — reliable aerial object acquisition remains the main bottleneck.

ACT
29.5%
Diffusion Policy
33.9%
π0
38.9%
π0.5
44.6%
Your method
?
??%

Mean success rate over all 16 Pick / Place × object combinations.

Can your policy beat π0.5?

Reliable aerial picking is still far from solved. Train and evaluate your method on the AeroManip-VLA dataset, and send us your results — we will add them to this leaderboard.

MethodSkillMaster Chef CanSugar BoxTomato Soup CanTuna Fish CanPudding BoxGelatin BoxPotted Meat CanBowlMean S.R.
ACTPick2.00.00.00.00.00.010.018.029.5
Place24.098.028.082.046.048.056.060.0
DPPick4.02.06.04.02.02.02.06.033.9
Place50.090.050.078.082.080.032.052.0
π0Pick6.76.710.020.013.313.313.36.738.9
Place80.082.052.088.052.068.086.024.0
π0.5Pick8.06.06.012.014.016.016.014.044.6
Place60.088.064.092.096.096.078.048.0

Success rate (%) per skill–object pair on AeroManip-VLA PrepareGroceries. Cell shading encodes the value; bold marks the best method per column and skill.

10 · Scalability

Parallelized GPU Simulation & Rendering

Each free-flying airframe is modeled with six virtual joints, so thousands of aerial manipulators — physics at 240 Hz, three onboard RGB-D cameras per step — run concurrently on one GPU.

2,137 SPS

with ~3,900 parallel environments using 32.7 GB

6.7×

throughput of AIR-VLA at comparable memory (317 SPS with 64 environments, 30.5 GB)

−86%

GPU memory vs. AIR-VLA at the same N = 64 (4.4 vs. 30.5 GB), with 1.6× throughput

3×

From one aerial manipulator to a grid of parallel apartments on a single GPU.

Throughput vs GPU memory
Simulation-and-rendering throughput vs. GPU memory on an RTX 5090 (log scale). Point labels give the number of parallel environments. AeroManip-VLA scales to ~3,900 environments (2,137 SPS at 32.7 GB), extending beyond AIR-VLA's last point (64 environments, 30.5 GB); at comparable memory it achieves about 6.7× the throughput.

BibTeX

@article{huang2026aeromanipvla,
  title   = {AeroManip-VLA: Scalable Vision-Language-Action Learning for
             Aerial Manipulation with RL-Generated Demonstrations},
  author  = {Huang, Rui and Mu, Yanlin and Li, Lidong and Wang, Yucong
             and Yan, Zichen and Zhao, Lin},
  journal = {arXiv preprint arXiv:2609.36915},
  year    = {2026}
}