ValueError: setting an array element with a sequence (inhomogeneous shape)
Your rows have different lengths, so NumPy cannot build a rectangular array from them. Find the odd row out, then pad the data or fix the source.
Updated
The error
ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (3,) + inhomogeneous part.
Older NumPy versions raise only the first sentence. The cause is identical.
What it means
A NumPy array is a rectangle (or box, or hyper-box): every row the same length, every layer the same shape. You handed np.array a nested list where the pieces have different sizes — ragged data. There is no rectangle that fits, so construction fails.
np.array([[1, 2, 3], [4, 5], [6, 7, 8]]) # row 2 is shortThe message decodes as: the outer dimension has 3 items, but their insides disagree.
Why it happens
Ragged data is usually honest — sentences have different lengths, patients have different numbers of visits, audio clips have different durations. The bug is expecting a rectangle from naturally uneven data.
Other routes: a parser that occasionally drops a field, mixing scalars and lists in one column, or a pandas column of lists reaching np.array via .values.
How to fix it
1. Find the odd rows.
lengths = [len(r) for r in rows]
print(set(lengths))
bad = [i for i, r in enumerate(rows) if len(r) != max(lengths)]
print(bad[:10])If one length dominates and a few rows differ, those rows are data bugs — inspect and repair them at the source.
2. If the data is naturally ragged, pad to a common length. This is the standard move before any model.
maxlen = max(len(r) for r in rows)
padded = np.zeros((len(rows), maxlen), dtype=np.float32)
for i, r in enumerate(rows):
padded[i, : len(r)] = rFor token sequences in PyTorch, torch.nn.utils.rnn.pad_sequence does this in one call. Hugging Face tokenizers do it with padding=True.
3. If you truly want an array of lists, say so explicitly.
arr = np.array(rows, dtype=object)Note what you give up: no vectorised maths, and most ML libraries will reject it downstream. This is for storage, not computation.
4. Check for accidental scalar/list mixing.
np.array([1, [2, 3]]) # same error, different bugOne element is a number, another is a list — usually a data-cleaning slip upstream.
How to prevent it
Validate lengths at the point where records are created or parsed, with a clear assert. Decide the padding policy once — length, pad value, truncation side — and apply it in one function. Keep ragged data in lists until the final padded conversion, so intermediate code makes no false rectangle promises.
Related errors
- stack expects each tensor to be equal size — the same problem inside a PyTorch DataLoader
- Object arrays cannot be loaded when allow_pickle=False — where dtype=object arrays cause trouble later
- operands could not be broadcast together