Job Data for
AI and ML Training.
Large-scale, AI-ready job data for training and evaluating models across employment, skills, occupations, and labour markets. Pre-enriched. LLM-friendly formats. MCP, S3, Parquet, or API delivery.
AI Models Are Only as Useful
as Their Training Data
Raw job postings contain the right signal, but inconsistent titles, buried skills, and missing salary make them hard to train on at scale. Propellum structures them before they reach your pipeline.
Raw job postings
Model-ready job data
AI models
Making job language
machine-readable.
Propellum applies a data quality layer designed to make large-scale job data usable for machine learning and AI development.
Build Models That Learn
From Market Change
Build the model around the question you want to answer, rather than relying on a predefined interpretation of the data.
Analytical Model Development
What the model can learn
Skills forecasting
Which skills are becoming more prevalent across roles and industries, and which are declining
Role classification
How job descriptions map to consistent occupations across varied employer title conventions
Demand prediction
Where hiring demand is increasing or declining by function, geography, and seniority
Salary modelling
Relationships between role, skills, location, seniority, and compensation at market scale
Job matching
Semantic similarity between job descriptions, skills profiles, and candidate experience
Trend detection
Emerging technologies, new role categories, and shifting skill clusters as they appear in employer postings
Skills forecasting
Which skills are becoming more prevalent across roles and industries, and which are declining
Role classification
How job descriptions map to consistent occupations across varied employer title conventions
Demand prediction
Where hiring demand is increasing or declining by function, geography, and seniority
Salary modelling
Relationships between role, skills, location, seniority, and compensation at market scale
Job matching
Semantic similarity between job descriptions, skills profiles, and candidate experience
Trend detection
Emerging technologies, new role categories, and shifting skill clusters as they appear in employer postings
Built for AI and data teams
that build with job data.
AI talent matching platforms
Training data for matching models, enriched job records with structured skills, normalised titles, and salary across 100+ countries.
Skills intelligence platforms
Ground truth for skills classification and extraction models. 40,000+ term taxonomy with required vs preferred labels per record.
Salary benchmarking tools
Training corpus for compensation estimation. 70% salary coverage with stated vs estimated flag, role, location, seniority, and skills as covariates.
Labour market intelligence platforms
28 years of structured data for temporal forecasting models. Role-level demand series across 100+ countries in a consistent schema.
LLM and foundation model teams
Diverse, multilingual job corpus for domain-specific pre-training or fine-tuning. 40+ languages, 100+ countries, 1998 to present.
Workforce analytics products
Historical and live job data as model input for skills demand forecasting, attrition risk, and workforce composition modelling.
The data to power
the models
| DIMENSION | SPECIFICATION |
|---|---|
| Volume | 5B+ deduplicated, enriched records |
| Historical depth | 1998 to present, 28 years |
| Geographic coverage | 100+ countries |
| Languages | 40+ languages in source text |
| Source | Employer career pages, not job boards or aggregators |
| Update frequency | Continuous, within hours of career page posting |
| Skills taxonomy | 40,000+ terms, required and preferred per record |
| Title normalisation | Proprietary taxonomy applied to all 5B+ records |
| Salary coverage | 70% of records, stated or AI-estimated with flag |
| Deduplication | Applied before delivery |
| Point-in-time | Available for temporal training splits |
| Delivery | MCP, S3, REST API, SFTP, data warehouse (Snowflake, BigQuery, Redshift) |
| AI-ready formats | JSONL (LLM fine-tuning), Parquet (Spark / pandas), JSON, CSV |
| Framework compat. | Compatible with PyTorch, TensorFlow, Hugging Face Datasets, LangChain, LlamaIndex |
Available via Model Context Protocol.
MCPPropellum job data is accessible via MCP — letting AI agents and LLM applications query structured job records directly as context, without a separate ETL pipeline.
AI & Machine Learning,
frequently asked questions
Build better models
with better job data.
Tell us your use case, target geographies, and preferred format. We send a structured sample in Parquet or JSON with full field definitions and coverage statistics within 24 hours.