Software Engineer - Storage & Embeddings
Bring native vector storage and semantic search into mosaicod (storage model, indexing, and query-engine integration), built in Rust and shipped open source.
Mosaico is an open-source data platform for robotics and physical AI. We’re solving a hard infrastructure problem: the tools that exist today for storing, indexing, and retrieving high-frequency sensor data are either too primitive or too painful to work with at scale.
We’re building the missing piece, and doing it in the open. Apache 2.0 licensed, public roadmap, real community. We care about the codebase being something people actually want to read and contribute to, not just use. In Italy, companies that operate this way are still a rarity, and we think that’s exactly what makes this interesting.
We’re based in Reggio Emilia, working with teams in defence, automotive, and agritech, and we’re building a commercial offering on top of the open-source core.
The role
Mosaico is expanding its platform to natively support embeddings. The goal is to allow users to enrich their sensor data with vector representations and run semantic search, similarity queries, and other embedding-based operations directly inside the platform, without having to move data out to an external system.
You’ll be the engineer who builds this layer. That means designing the storage model, the indexing strategy, the API surface that researchers and developers will use to interact with embeddings, and the integration with the existing query engine. There is significant design work ahead: how embeddings live alongside the existing data model, how they are indexed for performance at scale, and what the right abstractions look like for the people who will actually use them.
The work is in Rust, lives inside mosaicod, and is built in the open under Apache 2.0.
What you’ll work on
- Design the API surface that researchers and platform users will use to work with embeddings
- Define the storage model for vector data inside mosaicod and how it integrates with the existing data model
- Design and implement the indexing layer for high-performance similarity search at scale
- Integrate embedding-based search with the existing sequence, topic, and ontology query engine
- Contribute to the open-source codebase and participate in technical discussions with the community
What we’re looking for
- Deep experience with databases and storage systems, with a strong understanding of how they work under the hood
- Hands-on experience with vector databases or vector search libraries in a production context
- Experience designing APIs for technical users such as researchers and data engineers
- Strong Rust, or a demonstrated ability to get there fast coming from C++ or another systems language
- Solid understanding of indexing strategies and similarity search at scale
- Autonomous and opinionated: this role requires making hard architectural calls with a small team
- Fluent in English, written and spoken
Nice to have
- Familiarity with Apache Arrow or columnar storage formats
- Background in information retrieval or similarity search at scale
- Experience with embedding models and pipelines
- Open-source contributions in the data or systems space
What we offer
- Full design ownership of a layer that will define how Mosaico evolves toward physical AI use cases
- A small core team working on hard infrastructure problems in the open
- Work that ships as open source under Apache 2.0 and powers a commercial product
- Flexible hours, full remote with optional office access in Reggio Emilia
- Competitive salary, discussed openly based on your level and experience
- Welfare package
- Stock options
Salary range
€70K – 100K, adjusted based on location, experience, and level.
Contacts
Interested? Write to [email protected]