Given an instruction and observation, a generalist VLM first decomposes the task into a sequence of parameterized manipulation skills. At the start of each skill, the same model predicts a semantic target point in normalized image coordinates — the object to grasp, the handle to pull, or the placement location. A skill-specific trace expert then generates a 2D end-effector trace anchored to that target, and the corresponding action expert executes low-level actions from the trace-overlaid observation.
The current end-effector position is applied as an inpainting-style constraint that clamps the trace start, while the semantic target is injected through Adaptive RMSNorm conditioning. A lightweight MLP head predicts skill progress, advancing the pipeline to the next skill once a threshold is crossed. Crucially, the generalist VLM is never fine-tuned — guidance is accepted from any high-quality off-the-shelf model.