All posts
Engineering 7 min read

Semantic search over Robotics Data: yes, but keep your feet on the ground

Everyone wants to type "cyclist cutting in at dusk" and have their data lake answer. Embeddings make curation and failure triage much faster, but finding similar data is not the same as identifying a safety-critical scenario. Here's why we built hybrid search with pre-filtering, and what would make us change our minds.

Gabriele Lini CEO & Co-founder

Every robotics data platform is being asked the same question right now:

“When can I just type ‘cyclist cutting in at dusk’ and have my data lake actually answer?”

I get it. It’s a good question and you should get that feature. But you should also be suspicious of anyone selling it as the solution to everything.

I’ve been around since the Tin Man was considered advanced robotics, so when someone tells you embeddings will solve all your data problems, it’s worth asking what they’re actually good for.

Embeddings are useful. They can make curation and failure triage much faster. But finding similar data is not the same thing as identifying a safety-critical scenario, and neither is enough to support a safety case. There is no philosophical point here: if the data does not support the conclusion, the search result is not enough.

That distinction matters for us, too, and it is one of the reasons we are taking a different approach with Mosaico.

The problem is economic before it’s semantic

Let’s start with a reality check.

Robots generate a firehose of data. Multimodal, timestamped, largely unstructured. Cameras, LiDAR, radar, IMUs, joint states. The whole nine yards, all ending up in formats like rosbag2 and MCAP.

Here’s the thing: the events that make the whole dataset worth keeping are ridiculously rare.

The kind of rare where collecting useful data at scale becomes a logistics problem. Waymo and Tesla - two of the largest fleets out there - don’t bother with raw streaming at scale, they use trigger-based collection and aggressive filtering, because the economics of storing and indexing everything are simply not viable.

At that level of rarity, uniform sampling is essentially useless.

There is evidence behind this: studies in offline reinforcement learning have repeatedly shown that any form of intelligent data curation beats uniform sampling, hands down. And the best signals tend to be model-driven, not hand-crafted heuristics.

💡 Embeddings are a great way to do that curation, but they’re an accelerator, not an oracle.

Where embeddings actually shine

There are three workflows where semantic search really delivers.

Rare event retrieval, by text or by example

An engineer sees a weird frame, or a clip where the robot did something stupid. They say “find me more events like this” and the system does. Simple, powerful, useful.

Deduplication and coreset selection

This one is a direct, measurable money-saver.

If you’ve worked with large-scale datasets, you know that most of the data is redundant. Dropping near-duplicates can significantly reduce training data size with minimal performance loss, cutting GPU costs and training time. For example, NVIDIA Cosmos does this on video with large-scale clustering, but the principle is the same across domains.

Failure triage

This one’s worth calling out. The time it takes an engineer to find similar events to a failure goes from days to minutes. There’s no metric more telling than that.

Notice what’s not on this list, though: precise scenario mining.

The structured query still wears the crown

If you’ve been following the RefAV benchmark, you know what I’m talking about. They built 10,000 natural language queries over 1,000 Argoverse 2 logs. The headline result was a warning: naively repurposing off-the-shelf vision-language models doesn’t work well.

The most effective approaches today don’t rely on dense vector retrieval at all. They use LLMs that generate code to query structured object tracks. The vision-language model is only used for binary scene verification. The best HOTA-Temporal score? Still around 36. Coarse temporal localization works. But figuring out which agents and exactly when is largely unsolved.

Meanwhile, a structured query over perception outputs like:

“Ego speed above threshold AND tracked pedestrian within five meters AND hard-brake event”

…is exact, reproducible, and defensible.

Waymo’s own long-tail mining pipeline is rule-based plus MLLM over hand-defined categories. Not dense retrieval.

Why this matters

Because our customers live in the world of certifications. ISO 21448, ISO 26262, UNECE R157. For them, safety evidence has to be traceable, reproducible, and independent. A cosine similarity score has no provenance and no audit trail. Embedding retrieval helps you find the scenarios. But a documented, deterministic query is what goes into the safety case.

So: vector search is a discovery aid that feeds a traceable, rule-based pipeline. Not a replacement.

Semantic search

SAFETY ENGINEER

Safety engineer

Here’s what we actually built

Our choice is hybrid search with pre-filtering.

We first apply structured constraints: device, time range, ODD tags, topic, perception-derived quantities. Vector similarity then ranks the remaining candidates. Pure dense retrieval without metadata filters is almost always the wrong architecture for robotics logs.

And we’re not alone in this. Foxglove, Applied Intuition, Voxel51, and Scale are converging on the same design. When several mature products have arrived at the same architecture independently, I would take that seriously rather than inventing a different design just to stand out.

Where we differ is elsewhere: on-premises and air-gapped by default. Most of our customers can’t let their data leave the building. A cloud infrastructure is not possible, and this constraint forces us to choose where we compromise.

  • Above 500 million vectors, we use DiskANN indexes on NVMe.
  • Below 50 million, HNSW.
  • “Cold” vectors live on object storage with an SSD hot cache. It costs 100x less than RAM.

The price is a cold-query penalty in the hundreds of milliseconds versus single-digit warm. Fine for batch curation, painful for interactive triage unless you cache the hot partitions.

Better to know the trade-offs now than to learn about them the hard way.

Trade-offs

The one cost that keeps coming back

Embedding your data lake is a one-time cost: you do it and you move on. Re-embedding it is different.

As models improve, existing embeddings become stale. At large scale, re-embedding and reindexing can be expensive and time-consuming. There are drift adapter techniques that try to bridge the gap without full recomputation, and it’s a useful patch. But it’s still a patch.

For that reason, embedding versions should be part of the data model from the beginning. We model embeddings as versioned columns in existing tables, not a separate system of record you have to rebuild from scratch.

What would make us rethink this

This is our position today, not a commitment we are unwilling to revisit.

  • If scenario-mining benchmarks show dense retrieval crossing roughly 60 HOTA-Temporal, we’ll take a second look at dense-first architectures.
  • If a re-embedding-free adapter proves durable across two model generations, we’ll embed more aggressively.
  • If your situation allows data to leave the building, the storage-cost maths changes, and S3-backed managed engines are worth re-evaluating.

Until then, the lesson we’ve learned and wanted to share is: build your index behind your metadata, not instead of it.

TL;DR: for the impatient ones

TL;DR


Gabriele is CEO and co-founder of Mosaico. Since he stopped assembling IKEA furniture without reading the instructions, he dedicates himself to building data infrastructure that doesn’t make engineers cry. If you have a different take, he’d genuinely love to hear it.

Ready to tame
your data chaos?

Join disruptive companies and robotics leaders.
Start building with Mosaico in minutes.

Open source · Deploy in< 5 min