Description
Before an AI agent can reason about a contract, it needs trustworthy information to reason with.
At Monq, that information comes from complex enterprise procurement environments: legal documents, contracts, vendor records, pricing data, historical negotiations, ERP systems, live feeds, and internal business datasets. The Data Engineer builds the foundation that turns those sources into something AI systems can reliably use.
The work starts with ingestion. Pipelines need to handle high-volume contract and vendor data from multiple enterprise systems while accommodating documents and commercial structures that are far less standardized than ordinary application records.
Document parsing is therefore part of the challenge. Procurement contracts can contain prices, clauses, delivery requirements, service terms, risk conditions, timelines, and relationship factors buried inside legal language or semi-structured formats.
Once ingested, the data needs schemas and models capable of supporting contract intelligence, supplier benchmarking, negotiation analytics, and optimization. Discoverability and lineage matter because teams need to understand where a value came from and whether it can safely be used in a high-stakes commercial decision.
Data quality is especially important when autonomous agents consume the output. Missing, stale, duplicated, or incorrectly parsed information can influence a negotiation strategy, so completeness, latency, validation, audit trails, and governance are core engineering responsibilities.
The role also creates datasets and features for AI engineers. Real-time contract analysis, vendor research automation, and negotiation agents all depend on data infrastructure that can supply the right context without turning every new feature into a custom data project.
Technologies can include Python, distributed processing frameworks such as Spark or Flink, streaming systems such as Kafka, orchestration tools such as Airflow, Prefect, or Dagster, relational and NoSQL databases, data lakes, warehouses, APIs, queues, and cloud infrastructure.
Infrastructure-as-code, CI/CD, observability, incident response, performance tuning, and cost control are also relevant because these pipelines operate as production systems rather than offline research datasets.
This is an interesting opportunity for a data engineer who wants their work to sit directly underneath autonomous AI products rather than mainly powering dashboards or traditional business analytics.
Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.
https://careers.monq.io/jobs/7938369-data-engineer