AI & MACHINE LEARNING

Job Data for AI and ML Training.

Large-scale, AI-ready job data for training and evaluating models across employment, skills, occupations, and labour markets. Pre-enriched. LLM-friendly formats. MCP, S3, Parquet, or API delivery.

5B+
Pre-enriched records
40K+
Skills taxonomy
100+
Countries and languages
28+
Exp in job data
THE PROBLEM

AI Models Are Only as Useful as Their Training Data

Raw job postings contain the right signal, but inconsistent titles, buried skills, and missing salary make them hard to train on at scale. Propellum structures them before they reach your pipeline.

Stage 01

Raw job postings

messy terminology
duplicates
inconsistent descriptions
Stage 02

Model-ready job data

normalized
deduplicated
enriched
Stage 03

AI models

classification
matching
prediction
analysis
RAW VS ENRICHED

Making job language machine-readable.

Propellum applies a data quality layer designed to make large-scale job data usable for machine learning and AI development.

RAW, AS POSTED BY THE EMPLOYER
"title"
"Staff SDE IV (L6) - Hybrid"
ENRICHED, PROPELLUM OUTPUT
title_raw
"Staff SDE IV (L6) - Hybrid"
title_normalised
"Principal Software Engineer"
seniority
"Senior / Staff"
RAW, AS POSTED BY THE EMPLOYER
"description" (excerpt)
"...experience with distributed systems, strong Python skills, familiarity with Kubernetes and cloud infrastructure (AWS preferred), mentoring junior engineers..."
ENRICHED, PROPELLUM OUTPUT
skills_required
["Python", "Distributed Systems", "Kubernetes", "AWS"]
skills_preferred
["Team Leadership", "Mentoring"]
RAW, AS POSTED BY THE EMPLOYER
"salary"
"Competitive compensation package"
ENRICHED, PROPELLUM OUTPUT
salary_min
"$210,000"
salary_max
"$255,000"
salary_estimated
"true"
RAW, AS POSTED BY THE EMPLOYER
"location"
"Seattle, WA (Hybrid)"
ENRICHED, PROPELLUM OUTPUT
location_city
"Seattle"
location_country
"US"
coordinates
"47.6062, -122.3321"
remote_type
"Hybrid"
THE TAXONOMY LAYER

Build Models That Learn From Market Change

Build the model around the question you want to answer, rather than relying on a predefined interpretation of the data.

Analytical Model Development

STAGE 01
Historical Job Data
STAGE 02
Feature Extraction
STAGE 03
Pattern Recognition
STAGE 04
Model Training
STAGE 05
Prediction / Classification

What the model can learn

01

Skills forecasting

Which skills are becoming more prevalent across roles and industries, and which are declining

02

Role classification

How job descriptions map to consistent occupations across varied employer title conventions

03

Demand prediction

Where hiring demand is increasing or declining by function, geography, and seniority

04

Salary modelling

Relationships between role, skills, location, seniority, and compensation at market scale

05

Job matching

Semantic similarity between job descriptions, skills profiles, and candidate experience

06

Trend detection

Emerging technologies, new role categories, and shifting skill clusters as they appear in employer postings

BUILT FOR

Built for AI and data teams that build with job data.

AI talent matching platforms

Training data for matching models, enriched job records with structured skills, normalised titles, and salary across 100+ countries.

Skills intelligence platforms

Ground truth for skills classification and extraction models. 40,000+ term taxonomy with required vs preferred labels per record.

Salary benchmarking tools

Training corpus for compensation estimation. 70% salary coverage with stated vs estimated flag, role, location, seniority, and skills as covariates.

Labour market intelligence platforms

28 years of structured data for temporal forecasting models. Role-level demand series across 100+ countries in a consistent schema.

LLM and foundation model teams

Diverse, multilingual job corpus for domain-specific pre-training or fine-tuning. 40+ languages, 100+ countries, 1998 to present.

Workforce analytics products

Historical and live job data as model input for skills demand forecasting, attrition risk, and workforce composition modelling.

DATASET SPECIFICATION

The data to power the models

Volume
5B+ deduplicated, enriched records
Historical depth
1998 to present, 28 years
Geographic coverage
100+ countries
Languages
40+ languages in source text
Source
Employer career pages, not job boards or aggregators
Update frequency
Continuous, within hours of career page posting
Skills taxonomy
40,000+ terms, required and preferred per record
Title normalisation
Proprietary taxonomy applied to all 5B+ records
Salary coverage
70% of records, stated or AI-estimated with flag
Deduplication
Applied before delivery
Point-in-time
Available for temporal training splits
Delivery
MCP, S3, REST API, SFTP, data warehouse (Snowflake, BigQuery, Redshift)
AI-ready formats
JSONL (LLM fine-tuning), Parquet (Spark / pandas), JSON, CSV
Framework compat.
Compatible with PyTorch, TensorFlow, Hugging Face Datasets, LangChain, LlamaIndex

Available via Model Context Protocol.

MCP

Propellum job data is accessible via MCP — letting AI agents and LLM applications query structured job records directly as context, without a separate ETL pipeline.

Common Questions

AI & Machine Learning, frequently asked questions

FREE DATA SAMPLE

Build better models with better job data.

Tell us your use case, target geographies, and preferred format. We send a structured sample in Parquet or JSON with full field definitions and coverage statistics within 24 hours.