Robotics, Control and Autonomy
Vision-language-action models
A vision-language-action model reads a camera image and a plain-English instruction together, and outputs a robot action, using the same transformer ideas behind modern language models.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A vision-language-action model looks at a scene, reads an instruction in plain English, and outputs what the robot should do.
Think about asking a helper "hand me the towel on the chair." You did not point at exact coordinates. They looked around, worked out which object you meant, and picked it up. That combination — understanding words, looking at a scene, and acting — is exactly what a vision-language-action model is built to do. It is often shortened to VLA.
Why it exists
Older robot systems needed a separate, hand-built program for almost every task and every object. Teaching a robot to pick up a new kind of cup often meant new code, not only a new sentence of instruction.
VLA models take the transformer architecture behind modern language and vision-language models. That architecture is covered in transformers and vision transformers; a VLA model extends it one step further. Instead of only producing text, the model produces robot actions: gripper positions, joint movements, or similar low-level commands.
How it works
camera image + "pick up the red block and put it in the box"
|
v
[ VLA model: reads BOTH image and words together ]
|
v
robot action: move gripper toward (x, y, z) -> close gripper
-> move to the box -> open gripperTraining data for these models typically combines two sources. Large public image-and-text datasets — the same kind behind everyday vision-language AI — get mixed with robot-specific demonstration data, collected the way described in imitation learning.
Where you have already seen this idea
Research systems including Google DeepMind's RT-1 and RT-2, and the open-source OpenVLA project, follow this pattern. A household robot with a camera gets a spoken or typed instruction like "pick up the empty soda can." It identifies and picks up the correct object — even objects and phrasings it was never specifically programmed for.
What makes this different, and harder
A chatbot that gives a wrong answer produces an awkward sentence. A robot that misunderstands "pick up the cup" produces a real, physical, possibly damaging action in the real world. Language models can be wrong in ways that cost nothing. Robots acting on a misunderstanding cost something real — a broken object, at best.
An honest note
Vision-language-action models are an active, fast-moving research area, not a mature, production-proven technology. Published demonstrations are real, and genuinely impressive, and they are also typically shown under controlled conditions, with a limited set of objects and instructions. A model like this operating near people, or on anything valuable or fragile, needs far more validation than a research demo provides. That means extensive testing and human oversight. For many real applications, it also means a level of domain-expert and regulatory review that current VLA research has not yet gone through.
Remember this
- VLA models take an image and a plain-English instruction together, and output a robot action.
- They build directly on the transformer ideas behind modern vision and language AI.
- They remain research-stage technology; a demo working well is not the same as being ready for unsupervised real-world use.
What to learn next
- Safety in physical AI systems — what has to be true before a system like this is trusted near people.
- Transformers — the architecture VLA models are built on.
- Imitation learning and behaviour cloning — the main way these models learn what "pick up the cup" should physically look like.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
A drastically simplified stand-in for the VLA idea. An "instruction" — which colour to go to — and a "scene" — the positions of a red and a blue object — combine to predict the correct target position. There are no real images or real language here — that would need a model far too large for a laptop. Only the core mechanism remains: the same scene, a different instruction, a different action.
import numpy as np
from sklearn.neural_network import MLPRegressor
rng = np.random.default_rng(0)
def make_example():
red = rng.uniform(0, 10, size=2)
blue = rng.uniform(0, 10, size=2)
instruction = rng.choice(["go_to_red", "go_to_blue"])
instruction_onehot = [1, 0] if instruction == "go_to_red" else [0, 1]
target = red if instruction == "go_to_red" else blue
return [*instruction_onehot, *red, *blue], target
X_train, y_train = [], []
for _ in range(400):
features, target = make_example()
X_train.append(features)
y_train.append(target)
policy = MLPRegressor(hidden_layer_sizes=(32, 32), max_iter=3000, random_state=0)
policy.fit(X_train, y_train)
# A completely new object layout, never seen during training.
red_pos = np.array([2.0, 8.0])
blue_pos = np.array([7.0, 1.0])
go_red = policy.predict([[1, 0, *red_pos, *blue_pos]])[0]
go_blue = policy.predict([[0, 1, *red_pos, *blue_pos]])[0]
print(f"scene: red at {red_pos}, blue at {blue_pos} (never seen in training)")
print(f"instruction 'go to red' -> predicted target: [{go_red[0]:.2f}, {go_red[1]:.2f}] (true red: {red_pos})")
print(f"instruction 'go to blue' -> predicted target: [{go_blue[0]:.2f}, {go_blue[1]:.2f}] (true blue: {blue_pos})")scene: red at [2. 8.], blue at [7. 1.] (never seen in training) instruction 'go to red' -> predicted target: [2.39, 7.57] (true red: [2. 8.]) instruction 'go to blue' -> predicted target: [6.95, 0.90] (true blue: [7. 1.])
What actually happened
Both predictions used the exact same scene — only the instruction changed. The model correctly picked out a different target both times, close to the true position of the correct-coloured object. It never memorised fixed coordinates, because training used a fresh random layout every single example.
- The
instruction_onehotinput is the entire "language" side of this toy example. A real VLA model reads full sentences through a language model, instead of a two-number flag — but the underlying idea is the same: letting the instruction steer which part of the scene the output depends on. - Feeding both objects' positions in every input, regardless of which one the instruction asks for, forces the model to learn selection, not only memorise one fixed answer.
Common mistakes
Confusing this toy with a real VLA model. Real systems read raw camera pixels and free-form natural language, not two clean numbers and a one-hot flag. This example keeps the core "instruction selects the action" idea, and removes everything that requires a large pretrained model to work.
Assuming the model has learned "meaning." It learned a statistical mapping from a chosen input encoding to a target output, on a narrow synthetic task. This is a genuinely useful capability, and it is not the same claim as understanding language or objects the way a person does.
Testing only on layouts similar to training. The real value of a system like this is generalising to new objects, new phrasings, and new scenes it was never shown. Always test with a layout meaningfully different from anything in training, as this example does.
Try it yourself
Add a third colour, "go_to_green", with a third object position in the scene. Retrain, and check whether the model still correctly selects the right target for all three instructions. Watch how quickly a toy example like this needs real complexity once the task grows even slightly.
What to learn next
- Transformers — the real architecture behind the "vision" and "language" halves of an actual VLA model.
- Safety in physical AI systems — validation concerns that apply directly to any learned policy like this one.
- Imitation learning and behaviour cloning — how real systems learn the "pick up" and "put down" part of the action, not only the target position.
Researcher — Mathematics and papers.
Architecture
A VLA model typically extends a pretrained vision-language model (VLM) — one that already maps images and text to a shared representation — with an action decoding head. Two dominant formulations:
- Discretised action tokens. RT-2 (Brohan et al., 2023) represents each dimension of a continuous robot action (position, rotation, gripper state) as a discrete token from a fixed vocabulary, appended to the model's existing text vocabulary. This lets a single autoregressive transformer, already pretrained on web-scale image-text data, predict actions the same way it predicts text — using the exact same next-token-prediction objective. It is a direct application of ideas from decoding and generation in language models.
- Continuous action heads. Alternative architectures — Octo (Team et al., 2024) and OpenVLA (Kim et al., 2024) — attach a smaller diffusion or regression head atop a frozen or lightly fine-tuned VLM backbone. They predict continuous action chunks directly, rather than through discretisation.
Co-training on web and robot data
RT-2's central empirical claim is that co-training on internet-scale vision-language data alongside robot demonstration data improves generalisation to novel objects and instructions, compared to training on robot data alone. This is consistent with a broader pattern across the field. Capabilities learned from web-scale pretraining — recognising an unfamiliar object category, understanding an unusual phrasing — transfer into the robot action space, even though the web data itself contains no robot actions.
Cross-embodiment training
Different robots have different action spaces — a 6-DoF arm and a mobile base with a gripper are not interchangeable. The Open X-Embodiment dataset and RT-X models (Padalkar et al., 2023; Open X-Embodiment Collaboration, 2023) train a single policy across dozens of robot embodiments and labs simultaneously. Embodiment becomes another conditioning input, rather than a reason to train a fully separate model per robot. Positive transfer across embodiments is reported, though the degree varies significantly across specific robot pairs and tasks.
Evaluation
Standard robot learning benchmarks report task success rate under controlled conditions — a fixed, small library of objects and instructions, evaluated on real hardware or in simulation. This differs meaningfully from open-ended real-world deployment. Success rates on published benchmarks do not directly predict reliability across the much larger space of real objects, phrasings, and environments a genuinely general-purpose system would need to handle. No published VLA model reports evaluation at anything close to open-world, unconstrained deployment scale.
Current state and open problems
Inference latency remains a genuine constraint. Large VLA models can be too slow for control loops needing responses in tens of milliseconds. This motivates smaller distilled or hierarchical architectures — a large, slow model for high-level planning, paired with a small, fast model for low-level control. Safety and reliability guarantees for VLA-controlled physical actions are an active, unresolved research area. No current published system has demonstrated the kind of extensively validated reliability record that would support unsupervised deployment in safety-critical or unstructured public settings.
Key references
- Brohan, A. et al. (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. — arxiv.org/abs/2307.15818
- Padalkar, A. et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models.
- Kim, M. J. et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model.
- Team, Octo Model Team et al. (2024). Octo: An Open-Source Generalist Robot Policy.
What to learn next
- Safety in physical AI systems — the validation gap between benchmark performance and safe deployment.
- Transformers — the shared architecture behind the vision, language, and action pathways.
- Sim-to-real transfer — how the training data underlying many of these models is partly generated.