Description
Modern RAG quality depends on much more than embedding a document and returning the nearest vectors. Query planning, filtering, hybrid retrieval, indexing, ranking, metadata, knowledge structure, latency, and observability all influence what finally reaches the language model.
Pinecone is hiring at the Senior/Staff level to build that retrieval layer.
The engineer will work on components that connect structured and unstructured information to LLM applications and autonomous agents. Projects include semantic and hybrid search, metadata-aware retrieval, indexing pipelines, knowledge-graph construction, retrieval orchestration, APIs, and evaluation systems that measure whether search is actually returning useful context.
The underlying infrastructure operates at significant scale, so backend architecture and distributed-systems experience are important. Pinecone specifically looks for engineers who can reason about throughput, latency, reliability, security, and cost instead of optimizing retrieval quality in isolation.
Experience with vector search, Elasticsearch/OpenSearch-style systems, RAG pipelines, embeddings, query planning, retrieval evaluation, or knowledge systems maps particularly well to the role. Strong proficiency in languages such as Go, Rust, C++, Java, or Python is useful, along with Kubernetes, cloud infrastructure, observability, and infrastructure-as-code tooling.
There is also a strong agentic component: APIs and retrieval systems increasingly need to serve autonomous software consumers rather than only human-facing applications. Experience designing infrastructure for agent reasoning loops or multi-step retrieval is therefore valuable.
Pinecone currently lists this title in both US Remote and New York City variants. This entry uses the US Remote posting.
Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.
https://www.pinecone.io/careers/7ef089cb-a721-4ad8-a6d0-c390e64991d2/