Robotics, Control and Autonomy

Grasping and manipulation

Grasping is choosing where and how hard to close a gripper on an object, using its shape and the gripper's own physical limits.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest note
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Grasping is a robot working out where, and how hard, to close its gripper on an object.

You already know how to pick up a raw egg without cracking it, and a cricket ball without dropping it. Your fingers adjust before they even close, based on what your eyes tell you about the object. Grasping is teaching a robot that same skill: look at an object, decide where to place the gripper, and how firmly to squeeze.

Why it exists

Earlier lessons in this platform cover looking at the world — image classification, object detection. Grasping is the next step: touching it, and picking it up, without dropping it or crushing it.

This turns out to be much harder than recognising the object. A coffee mug and a bar of soap look nothing alike. Yet both need the gripper to close somewhere specific, at a specific angle, with a specific force. Not too little — it slips. Not too much — it deforms or breaks.

How it works

camera / depth sensor  -->  see the object's shape and position
        |
        v
candidate grasp points  -->  where could the gripper close?
        |
        v
score each candidate   -->  how well do the gripper's two jaws
                             oppose each other here? does the
                             gripper even open wide enough?
        |
        v
best grasp  -->  move gripper there  -->  close  -->  check for slip

A common, useful idea is the antipodal grasp. Pick two points on the object where the surface faces almost directly opposite each other. Then the gripper's two jaws squeeze straight toward one another, instead of at an angle — and that is far less likely to slip.

Where you have already seen it

  • Warehouse robots that pick individual items out of a bin of mixed products, a task the industry calls "bin picking."
  • Fruit-picking robots being trialled in orchards, which must judge ripeness and stem attachment before gripping.
  • Robotic dishwashing and sorting demonstrations, which struggle with anything reflective, transparent, or oddly shaped — for reasons covered below.

An honest note

General-purpose robot grasping — picking up almost anything, reliably, the way a person can — remains an open research problem, not a solved one. Transparent glass, reflective metal, and soft or deformable objects still cause real failures in current systems. A robot that grasps well in a lab demo is not the same as one certified to work safely around people. Assistive or care robots that would touch a person's body need far more rigorous, domain-expert-validated testing than anything shown in this lesson.

Remember this

  • Grasping means choosing where to place the gripper and how hard to close it, from what a sensor sees.
  • An antipodal grasp — jaws facing straight at each other — resists slipping better than an angled one.
  • The gripper's own physical limits, like maximum opening width, constrain which grasps are even possible.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

A simplified antipodal grasp search. The object is an ellipse — long and thin, like a TV remote lying on a table. It is represented as boundary points with outward-facing normals, the kind of shape a depth camera and a segmentation model would hand you.

antipodal_grasp.py
import numpy as np

theta = np.linspace(0, 2 * np.pi, 60, endpoint=False)
a, b = 0.12, 0.03   # semi-major, semi-minor axis in metres -- 24cm long, 6cm wide
points = np.stack([a * np.cos(theta), b * np.sin(theta)], axis=1)

normals = np.stack([np.cos(theta) / a, np.sin(theta) / b], axis=1)
normals /= np.linalg.norm(normals, axis=1, keepdims=True)

GRIPPER_MAX_WIDTH = 0.09  # this gripper opens 9cm -- too narrow for the 24cm length

def score_pair(i, j):
    p1, p2, n1, n2 = points[i], points[j], normals[i], normals[j]
    width = np.linalg.norm(p1 - p2)
    if width > GRIPPER_MAX_WIDTH or width < 0.01:
        return None
    alignment = -np.dot(n1, n2)   # 1.0 = normals point straight at each other
    return alignment - 0.5 * width

best_score, best_pair = -np.inf, None
for i in range(len(points)):
    for j in range(i + 1, len(points)):
        result = score_pair(i, j)
        if result is not None and result > best_score:
            best_score, best_pair = result, (i, j)

i, j = best_pair
p1, p2 = points[i], points[j]
print(f"object: {2*a*100:.0f}cm long x {2*b*100:.0f}cm wide, gripper opens {GRIPPER_MAX_WIDTH*100:.0f}cm")
print(f"best grasp width found: {np.linalg.norm(p1 - p2)*100:.1f}cm")
print(f"grasp points (metres): {p1.round(3)} and {p2.round(3)}")
print("this grasps across the object's SHORT dimension -- the only one the gripper can close on")
Output
object: 24cm long x 6cm wide, gripper opens 9cm
best grasp width found: 6.0cm
grasp points (metres): [0.   0.03] and [-0.   -0.03]
this grasps across the object's SHORT dimension -- the only one the gripper can close on

What actually happened

The search checked every pair of boundary points. It threw out any pair too wide for the gripper, or too close together to be useful. What remained was the pair whose surface normals pointed most directly at each other. It correctly found the object's 6cm width, and never considered its 24cm length — because the gripper physically cannot open that far.

  • alignment = -np.dot(n1, n2) is 1.0 when two normals point exactly opposite each other, and less for any angle away from that — this is what "antipodal" means numerically.
  • if width > GRIPPER_MAX_WIDTH: return None is doing real work here. Remove it, and the search would happily suggest a grasp the gripper cannot physically perform.

Common mistakes

Ignoring the gripper's own limits during search. A grasp that looks geometrically ideal but exceeds the gripper's maximum width, minimum width, or reach is not a candidate at all. Filter it out before scoring, not after.

Scoring only geometry, never force. This example never asks how heavy the object is, or how much force is needed to lift it without slipping. Real grasp planners incorporate an estimate of friction and mass, often refined using force sensors in the fingers after contact.

Assuming a 2D outline is enough. A real object is 3D. A grasp that looks fine from one camera angle can be blocked by the table from another — or miss the object's centre of mass entirely. Production systems reason in 3D, usually from depth or multiple camera views.

Try it yourself

Change GRIPPER_MAX_WIDTH to 0.30 — a gripper that can open wider than the object is long. Rerun the search and see which grasp width wins this time, and why grasping the long dimension can be a worse choice even when it is possible.

What to learn next

Researcher — Mathematics and papers.

The grasp quality problem

Given an object's surface (a mesh, point cloud, or signed distance field) and a gripper model, grasp planning searches over a pose for the gripper — 6-DoF, or more for multi-fingered hands — plus finger positions. The goal is maximising an expected grasp quality metric, subject to reachability and collision constraints.

The classical force-closure criterion asks whether the set of contact forces the gripper can apply is capable of resisting an arbitrary external wrench (force and torque) applied to the object. Ferrari and Canny (1992) formalise a widely used quality metric: the radius of the largest ball, centred at the origin, contained within the convex hull of achievable contact wrenches. A larger radius means the grasp resists disturbance in more directions, by more force, before failing.

Learned grasp detection

Modern approaches skip explicit force-closure computation and instead learn to predict grasp quality directly from raw sensor data. Two influential lines of work:

  • Dex-Net (Mahler et al., 2017) trains a CNN on millions of synthetically generated grasp attempts, simulated against a database of 3D object models with analytic force-closure labels, then transfers to real depth-camera images.
  • GG-CNN / grasp-quality CNNs predict a per-pixel grasp quality, angle, and width map directly from a single depth image, enabling closed-loop grasp planning at video frame rate rather than a separate offline planning step.

Tactile and closed-loop grasping

Vision alone cannot resolve friction, exact contact point, or slip. Tactile sensors (GelSight, BioTac, and similar) provide direct contact information after the gripper closes. Slip detection — noticing an object beginning to move relative to the gripper — allows closed-loop force adjustment: increase grip force reactively, rather than relying on an open-loop force estimate chosen before contact.

Deformable and articulated objects

Rigid-body grasp theory does not directly apply to cloth, cables, or food items that deform under grasp force. This remains an active research area, generally requiring either learned deformation models, or continuous visual and tactile feedback during the grasp rather than a single pre-computed grasp pose.

Current state

Bin-picking of rigid, opaque, previously-seen objects is largely a solved engineering problem in industrial settings today. Grasping novel, transparent, reflective, or deformable objects reliably, across arbitrary environments, remains unsolved. It is an active benchmark area — the annual "Amazon Picking Challenge" successor competitions, and academic benchmarks like the YCB object set, among others.

Key references

  • Ferrari, C., Canny, J. (1992). Planning Optimal Grasps. ICRA.
  • Mahler, J. et al. (2017). Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics. RSS.
  • Morrison, D., Corke, P., Leitner, J. (2018). Closing the Loop for Robotic Grasping: A Real-time, Generative Grasp Synthesis Approach. RSS — GG-CNN.
  • Calandra, R. et al. (2018). More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch. IEEE Robotics and Automation Letters.

What to learn next