Job

Data Engineer — Data Infrastructure for Autonomous AI Agents

Trainety Curated Opportunities

Location
Poland
Industry
Technology & Internet
Organization size
Individual
Updated
September 12, 2026

Description

Before an AI agent can reason about a contract, it needs trustworthy information to reason with.


At Monq, that information comes from complex enterprise procurement environments: legal documents, contracts, vendor records, pricing data, historical negotiations, ERP systems, live feeds, and internal business datasets. The Data Engineer builds the foundation that turns those sources into something AI systems can reliably use.

The work starts with ingestion. Pipelines need to handle high-volume contract and vendor data from multiple enterprise systems while accommodating documents and commercial structures that are far less standardized than ordinary application records.

Document parsing is therefore part of the challenge. Procurement contracts can contain prices, clauses, delivery requirements, service terms, risk conditions, timelines, and relationship factors buried inside legal language or semi-structured formats.

Once ingested, the data needs schemas and models capable of supporting contract intelligence, supplier benchmarking, negotiation analytics, and optimization. Discoverability and lineage matter because teams need to understand where a value came from and whether it can safely be used in a high-stakes commercial decision.

Data quality is especially important when autonomous agents consume the output. Missing, stale, duplicated, or incorrectly parsed information can influence a negotiation strategy, so completeness, latency, validation, audit trails, and governance are core engineering responsibilities.

The role also creates datasets and features for AI engineers. Real-time contract analysis, vendor research automation, and negotiation agents all depend on data infrastructure that can supply the right context without turning every new feature into a custom data project.

Technologies can include Python, distributed processing frameworks such as Spark or Flink, streaming systems such as Kafka, orchestration tools such as Airflow, Prefect, or Dagster, relational and NoSQL databases, data lakes, warehouses, APIs, queues, and cloud infrastructure.

Infrastructure-as-code, CI/CD, observability, incident response, performance tuning, and cost control are also relevant because these pipelines operate as production systems rather than offline research datasets.

This is an interesting opportunity for a data engineer who wants their work to sit directly underneath autonomous AI products rather than mainly powering dashboards or traditional business analytics.

Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.

https://careers.monq.io/jobs/7938369-data-engineer

Expertise

  • Data Engineering
  • Python
  • ETL Pipelines
  • Streaming Data
  • Kafka
  • Data Quality
  • Terraform
  • Curated Opportunity

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.