Error database

ValueError: setting an array element with a sequence (inhomogeneous shape)

Your rows have different lengths, so NumPy cannot build a rectangular array from them. Find the odd row out, then pad the data or fix the source.

The message you saw
ValueError: setting an array element with a sequence (inhomogeneous shape)

By Updated

The error

Output
ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (3,) + inhomogeneous part.

Older NumPy versions raise only the first sentence. The cause is identical.

What it means

A NumPy array is a rectangle (or box, or hyper-box): every row the same length, every layer the same shape. You handed np.array a nested list where the pieces have different sizes — ragged data. There is no rectangle that fits, so construction fails.

python
np.array([[1, 2, 3], [4, 5], [6, 7, 8]])    # row 2 is short

The message decodes as: the outer dimension has 3 items, but their insides disagree.

Why it happens

Ragged data is usually honest — sentences have different lengths, patients have different numbers of visits, audio clips have different durations. The bug is expecting a rectangle from naturally uneven data.

Other routes: a parser that occasionally drops a field, mixing scalars and lists in one column, or a pandas column of lists reaching np.array via .values.

How to fix it

1. Find the odd rows.

python
lengths = [len(r) for r in rows]
print(set(lengths))
bad = [i for i, r in enumerate(rows) if len(r) != max(lengths)]
print(bad[:10])

If one length dominates and a few rows differ, those rows are data bugs — inspect and repair them at the source.

2. If the data is naturally ragged, pad to a common length. This is the standard move before any model.

python
maxlen = max(len(r) for r in rows)
padded = np.zeros((len(rows), maxlen), dtype=np.float32)
for i, r in enumerate(rows):
    padded[i, : len(r)] = r

For token sequences in PyTorch, torch.nn.utils.rnn.pad_sequence does this in one call. Hugging Face tokenizers do it with padding=True.

3. If you truly want an array of lists, say so explicitly.

python
arr = np.array(rows, dtype=object)

Note what you give up: no vectorised maths, and most ML libraries will reject it downstream. This is for storage, not computation.

4. Check for accidental scalar/list mixing.

python
np.array([1, [2, 3]])      # same error, different bug

One element is a number, another is a list — usually a data-cleaning slip upstream.

How to prevent it

Validate lengths at the point where records are created or parsed, with a clear assert. Decide the padding policy once — length, pad value, truncation side — and apply it in one function. Keep ragged data in lists until the final padded conversion, so intermediate code makes no false rectangle promises.

The lessons behind this error.

  • Python for AI

    NumPy

    NumPy lets you do one operation to millions of numbers at once instead of one at a time. It is the foundation every AI library in Python is built on.

Back to all errors