GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation

Ninghan Zhong1*†, Jing-Chen Peng1*, Sriram Vishwanath1
1Georgia Institute of Technology
*Equal contribution  ·  †Corresponding author

Code link is coming soon.

Overview video: method walkthrough, simulation and real-robot results.


Abstract

Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings.


Method

GT-VLA overview: a generalist VLM decomposes a task into skills and supplies a semantic target, which becomes an image-space trace conditioning the action policy.

Overview. An off-the-shelf generalist VLM decomposes tasks into skills and provides skill-specific guidance. This is converted into an image-space trace conditioning for the action policy, ensuring execution follows high-level intent.

Given an instruction and observation, a generalist VLM first decomposes the task into a sequence of parameterized manipulation skills. At the start of each skill, the same model predicts a semantic target point in normalized image coordinates — the object to grasp, the handle to pull, or the placement location. A skill-specific trace expert then generates a 2D end-effector trace anchored to that target, and the corresponding action expert executes low-level actions from the trace-overlaid observation.

The current end-effector position is applied as an inpainting-style constraint that clamps the trace start, while the semantic target is injected through Adaptive RMSNorm conditioning. A lightweight MLP head predicts skill progress, advancing the pipeline to the next skill once a threshold is crossed. Crucially, the generalist VLM is never fine-tuned — guidance is accepted from any high-quality off-the-shelf model.

GT-VLA architecture: a skill-routed trace expert generates a 2D trace, and a skill-routed action expert predicts actions from the trace-overlaid observation.

Trace-Action Mixture-of-Experts (TA-MoE). Left: a skill-routed trace expert generates a 2D trace anchored to the semantic target and the end-effector position. Right: a skill-routed action expert predicts actions from the trace-overlaid observation, and an MLP predicts skill progress. A deterministic router selects experts for the current skill.


Robust Trace Conditioning

A trace is only useful if the policy actually follows it — and only safe if the policy can ignore it when 2D guidance is geometrically ambiguous. Three augmentations balance these:

  • Random scene drop. The scene is masked while the rendered trace stays visible, forcing the policy to sometimes act from guidance alone and reducing overfitting to memorized training scenes.
  • Random trace drop. The trace overlay is removed, preserving the policy's ability to reason from the raw observation when depth, occlusion, or contact geometry matters.
  • Low-frequency trace perturbation. Training traces are bent by smooth noise, closing the gap between ground-truth traces at training time and generated traces at inference time.
Trace-conditioning augmentations shown on LIBERO and the real robot: original, perturbed trace, scene drop, and trace drop.

Augmentations shown on LIBERO (top row) and the real robot (bottom row). Removing any one of the three costs 10.5–13.2 percentage points of average success, and the trace interface yields limited benefit unless all three are present — they are complementary rather than redundant.


Real-World Results

A YAM arm with side-view and wrist cameras, trained on 299 episodes of long-horizon pick-and-place. Five evaluation suites: two composition suites (familiar skills in unseen orders, 4 and 6 skills) and three OOD suites (unseen skill configurations, 2, 4 and 6 skills). Each suite is 10 tasks × 5 trials; a trial counts only if the full task completes.

Real-world success rates across five evaluation suites, comparing GT-VLA against AtomicVLA and pi-0.5.

Real-world results, 5 trials × 10 tasks per category. GT-VLA outperforms both baselines across all five suites, and its advantage widens as the task horizon grows.

Rollouts

The clips below are individual trials, chosen to illustrate the aggregate result above rather than to stand in for it. GT-VLA succeeds in each; the baselines remain capable of pick-and-place motion but drift away from the task semantics under unseen instructions and configurations.

“Place both the red bell pepper and the corn into the basket, and place the carrot onto the blue plate.”

Real-world · 3 skills

GT-VLA (ours)

Success

Completes all three placements.

AtomicVLA

Failure

Places the red bell pepper on the blue plate, then the carrot in the basket — both wrong locations.

π0.5

Failure

Places the red bell pepper on the blue plate instead of in the basket.

“Put the orange plate onto the blue plate, and place both the asparagus and the corn on the orange plate.”

Real-world · 3 skills

GT-VLA (ours)

Success

Stacks the plates, then places both vegetables correctly.

AtomicVLA

Failure

Places the asparagus and corn on the blue plate, and skips the plate-stacking step entirely.

π0.5

Failure

Grasps the blue plate rather than the orange one.

“Move the green bell pepper onto the green plate, and place both the carrot and the asparagus onto the white plate.”

Real-world · 3 skills

GT-VLA (ours)

Success

Routes each item to its correct plate.

AtomicVLA

Failure

Places the carrot on the green plate instead of the white one.

π0.5

Failure

Places the green bell pepper on the white plate, and ignores the asparagus instruction.

The baselines can still perform pick-and-place motions, but struggle to stay aligned with the task semantics under unseen instructions and configurations. GT-VLA uses semantic target-conditioned traces to keep each skill grounded to the correct object and placement location.

LIBERO Results

To test generalization rather than in-distribution fitting, we forgo per-suite finetuning: all methods train on LIBERO-10 and LIBERO-90 and are evaluated on the held-out Goal, Spatial, and Object suites.

ModelGoalSpatialObjectAvg.
π0.516.2 (13.1–19.7)29.6 (25.6–33.8)1.6 (0.7–3.1)15.8
AtomicVLA30.2 (26.2–34.4)46.8 (42.4–51.3)24.6 (20.9–28.6)33.9
SEAL31.0 (27.0–35.3)45.0 (40.6–49.5)6.0 (4.1–8.5)27.3
GT-VLA (ours)38.8 (34.5–43.2)50.8 (46.3–55.3)41.8 (37.4–46.3)43.8

Success rate on LIBERO out-of-distribution evaluation suites, 50 trials × 10 tasks per category (%). 95% Wilson confidence intervals in parentheses.

GT-VLA leads every suite, averaging 43.8% — +28.0 points over π0.5, +9.9 over AtomicVLA, and +16.5 over SEAL. Both SEAL and GT-VLA consume guidance from an external generalist VLM, so the gap between them indicates that external guidance alone is not enough: the VLA must be designed to accept it.

Ablations

G-VLA removes the trace stage but keeps semantic-target guidance, injected directly into the action heads. GT-VLA-single removes the skill-routed MoE, using one trace head and one action head. The remaining rows are leave-one-out augmentation ablations and test-time swaps of the guidance VLM.

ModelGoalSpatialObjectAvg.
Architecture
GT-VLA-single27.6 (23.7–31.7)40.6 (36.3–45.0)27.8 (23.9–32.0)32.0
G-VLA32.2 (28.1–36.5)39.6 (35.3–44.0)20.8 (17.3–24.6)30.9
Augmentations (leave-one-out)
No scene drop38.4 (34.2–42.7)29.0 (25.2–33.1)24.4 (20.8–28.4)30.6
No trace drop36.6 (32.5–40.9)41.8 (37.6–46.2)21.4 (18.0–25.2)33.3
No perturbation34.8 (30.8–39.1)40.6 (36.4–45.0)23.8 (20.3–27.7)33.1
Guidance VLM swap†
GPT-5.6 Terra33.0 (26.9–39.8)44.5 (37.8–51.4)35.0 (28.7–41.8)37.5
Qwen 3.8 Max35.0 (28.7–41.8)37.0 (30.6–43.9)30.5 (24.5–37.2)34.2
Gemini 3.6 Flash31.0 (25.0–37.7)46.5 (39.7–53.4)33.5 (27.3–40.3)37.0
Full model
GT-VLA (ours)38.8 (34.5–43.2)50.8 (46.3–55.3)41.8 (37.4–46.3)43.8

Ablation results on LIBERO (%). 50 trials × 10 tasks per category, except the guidance VLM swap rows†, which use 20 trials × 10 tasks. 95% Wilson confidence intervals in parentheses. GT-VLA's default guidance model is Gemini 3.1 Pro.

Why the trace layer matters

Aggregate numbers say the trace stage and the MoE both help — 43.8% for the full model against 30.9% for G-VLA and 32.0% for GT-VLA-single. These LIBERO rollouts show the mechanism behind those numbers.

“Pick up the tomato sauce and place it in the basket.”

LIBERO-Object · out-of-distribution

GT-VLA (ours)

Success

Grasps the tomato sauce and completes the placement.

G-VLA (no trace)

Failure

Despite a target point on the tomato sauce, the policy picks up the BBQ sauce.

GT-VLA-single (no MoE)

Failure

Despite both the trace and the target point on the tomato sauce, the policy picks up the BBQ sauce.

“Pick up the black bowl on the stove and place it on the plate.”

LIBERO-Goal · out-of-distribution

GT-VLA (ours)

Success

Moves the bowl from the stove to the plate.

G-VLA (no trace)

Failure

Despite a target point on the bowl, the policy attempts to pick up the plate.

GT-VLA-single (no MoE)

Failure

Lifts the bowl off the stove, then immediately puts it back — confusing pick-from-stove with place-onto-stove.

GT-VLA, with MoE trace generation, is generally steerable and shows stronger generalization. G-VLA frequently fails despite receiving the same semantic guidance, suggesting that directly injecting a semantic target is insufficient in comparison to an actionable trace. GT-VLA-single also struggles to generalize even with trace guidance: without skill-specific experts, it can confuse skills that share similar trajectories, such as pick and place.

Robustness to Noisy Guidance

GT-VLA depends on a generalist VLM for semantic targets, so we perturb those targets deliberately: for a noise radius r, a new target is sampled uniformly from a disk of that radius around the original point. Observations are 224×224, so r = 64 spans over a quarter of the image width.

GT-VLA success rates across LIBERO suites under increasing semantic target noise radii.
Visual illustration of semantic target perturbations at increasing noise radii.

Semantic target perturbations. Top: success rates across LIBERO suites under noisy targets. Bottom: pink cross marks the given point, the white circle the maximum perturbation radius, and the red dot the perturbed final point.

Performance holds up to r = 32 px — where GT-VLA still generally beats the unperturbed baselines from the table above — because a target that lands on or beside the intended object still points the trace in a broadly correct direction. The sharp drop at r = 64 px marks the boundary: targets there routinely land on the wrong object or on empty background.

Failure Analysis

Each failed episode is assigned a single primary failure source — the earliest stage in the pipeline that deviated.

Failure breakdowns for GT-VLA across simulation, hardware, and noise-injection settings.

GT-VLA failure breakdowns for simulation, hardware, and noise injection.

The breakdown supports three design choices:

  • An off-the-shelf generalist suffices for sparse high-level guidance without fine-tuning — high-level VLM and progress-prediction errors stay a minority source across categories.
  • The trace-conditioned policy is reliably steerable — failures from not following the guidance trace range from just 0.6% to 8.3%, in sharp contrast to G-VLA, which executes memorized motions despite correct targets.
  • Errors do not strongly cascade — perturbing the semantic target, or extending real-world tasks from four to six skills, leaves the failure profile largely unchanged.

The remaining failures point to concrete next steps. In LIBERO, inaccurate traces dominate (36.5%). On hardware, low-level execution failures such as grasp slips lead (41.7% for 4-skill tasks) — these occur despite correct guidance and would benefit from stronger low-level controllers. And 2D–3D ambiguity consistently accounts for around 20%, a limitation of image-space guidance itself that motivates multi-view or 3D traces.


BibTeX

@article{zhong2026gtvla,
  title   = {GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation},
  author  = {Zhong, Ninghan and Peng, Jing-Chen and Vishwanath, Sriram},
  journal = {arXiv preprint arXiv:2609.31904},
  year    = {2026}
}