logoPandorLabs
job postings data

Hiring Data, Before the Announcement

Job postings from company career pages and the major boards, parsed into title, seniority, function, location, salary, and skills — deduplicated across sources, so one role posted in five places counts once.

at a glance

Product
Job Postings Data
Formats
JSON · CSV · Parquet
Delivery
Webhooks, S3, Snowflake, BigQuery
Refresh
Continuous to daily
Scope
Public data only

the problem

Why this is harder than it looks

A job posting is the most honest public document a company produces. Marketing describes what a company wants to be understood as doing; a requisition describes what it is actually paying people to do, which is a different and much more reliable thing. Read in aggregate, postings are a leading indicator that runs months ahead of announcements — a team is hired before a product ships, an office is staffed before a market entry is public, and a hiring freeze shows up in the requisition count well before it shows up in a filing. The difficulty is not finding postings but making them comparable: the same role appears on the career page and four boards under four titles, and "Senior Engineer II" means something different at every company.

who buys it

Who this is built for

Competitive intelligence teams reading hiring as a roadmap, investors and researchers using headcount direction as a fundamental signal, recruiters mapping where talent demand is concentrated, and HR teams benchmarking their own compensation and titling against a real market rather than a survey.

sources covered

What we collect from

Company career pages

The primary source, and the one that matters most. A posting appears here first and is often never syndicated at all for senior roles.

Applicant tracking systems

Greenhouse, Lever, Workday, Ashby and similar. Structurally consistent, which makes them the highest-fidelity source of posting data available.

Major job boards

General and regional boards, used to catch roles that never reach a public career page and to establish which listings are duplicates.

Specialist and industry boards

Niche boards where technical, clinical, and academic roles are posted and which general aggregators systematically miss.

Regional and non-English sources

Local-language boards and career pages, so a multinational’s hiring outside its home market is not invisible in your data.

Historical posting archives

Backfilled postings including ones since removed, since a role that was posted and then withdrawn is itself a meaningful signal.

what you get back

Fields in the delivered schema

Agreed with you before collection starts, and held stable afterwards — the sites change underneath, your columns do not.

  • Job title as posted, plus a normalised title for cross-company comparison
  • Seniority band and job function, mapped to a consistent taxonomy
  • Company name, resolved to a stable company identifier across sources
  • Location, with remote, hybrid, and on-site distinguished explicitly
  • Salary range where disclosed, with currency and period normalised
  • Full job description text
  • Extracted skills, tools, and technologies mentioned
  • First-seen and last-seen dates, giving time-to-fill and removal signals
  • Employment type — permanent, contract, internship — and department where stated
  • Deduplication cluster ID grouping the same role across every source it appeared on

Public surfaces only

We collect what a visitor can see, honour a site's stated crawling preferences, and never bypass authentication. Provenance is recorded on every record.

One schema across sources

Records from any source arrive with the same field names, so adding a source does not mean rewriting anything downstream.

Compliance built in

GDPR and CCPA handling, a DPA signed before delivery, configurable retention, and deletion at source propagating through to your feed.

Daily
Refresh cadence
Deduped
Across all sources
First+last
Seen dates
99.9%
Uptime SLA

applications

What teams build with job postings data

Competitor roadmap inference

What a company hires for is what it is building, months before anything ships. A run of infrastructure roles, a first compliance hire, or a sudden cluster of sales engineers each describe a strategy more literally than any announcement will.

Headcount direction as a fundamental

Posting volume by function and its trend is one of the earliest public reads on whether a company is expanding or contracting, and it moves well ahead of reported financials.

Market entry and expansion detection

Roles appearing in a new city or country are usually the first public evidence of an expansion, and they typically precede the press release by a quarter or more.

Compensation benchmarking

Disclosed salary ranges, normalised by currency, period, and level, give a picture of the real market rather than a survey with a twelve-month lag.

Skills and technology adoption tracking

Which tools and frameworks are being hired for, across a whole sector and over time. For anyone selling to engineering teams, this is a direct read on adoption.

Talent market mapping

Where demand for a given skill is concentrated geographically and which companies are competing for the same candidates you are.

related products

Usually bought alongside

Everything below delivers on the same schema and the same infrastructure, so combining them is a configuration change rather than a project.

questions

Frequently Asked Questions

Everything you need to know before you send us your first request.

Postings are clustered by company, normalised title, location, and description similarity, and every duplicate carries a shared cluster ID with the earliest sighting marked as the original. This matters more than it sounds: a company that syndicates aggressively can look like it is hiring five times as fast as one that only posts to its career page, which turns any un-deduplicated comparison between them into nonsense.
Yes, and it is the field that makes the dataset comparable at all. The title as posted is preserved verbatim, and alongside it we deliver a normalised title, a seniority band, and a function drawn from a consistent taxonomy. Without that layer, "Staff Engineer", "Senior Engineer II", and "Engineer III" are three unrelated strings rather than three points on one scale.
Where it is disclosed in the posting, yes — normalised for currency and period so an hourly rate and an annual band can sit in the same column. Disclosure rates vary enormously by jurisdiction, and in markets with pay transparency legislation coverage is high while elsewhere it is sparse. The field is explicitly null when nothing was disclosed rather than estimated, because an inferred salary presented as an observed one is worse than no data.
Yes. Every posting carries first-seen and last-seen dates, so a removal is visible as a last-seen date that stops advancing. That gives you time-to-fill as a derived metric and makes withdrawn roles detectable — a posting that appears and disappears within a week is itself a signal worth having.
Any company with a public career page, which in practice is nearly all of them. Engagements are usually scoped to a named watchlist — your competitors, your customers, a sector, or an index — rather than everything, because a targeted list at daily cadence is more useful than a firehose nobody filters.
No. Job postings describe roles rather than people, which makes this one of the cleanest datasets we deliver from a privacy standpoint. Where a posting names a hiring manager or recruiter, that is handled under the same lawful-basis and DPA framework as the rest of our contact data, and it can be excluded from delivery entirely if you would rather not receive it.
Into the systems you already run. Records land as JSON, CSV, or Parquet in S3, GCS, or Azure Blob, or straight into Snowflake or BigQuery, on whatever cadence you set. Webhooks push new records to your endpoints as they are detected, which is how most customers wire alerting. A solutions engineer sets the schema, cadence, and destination with you during onboarding rather than handing you a docs page and wishing you luck.
Keeping the collectors working is our responsibility, not yours. Extractors are monitored continuously, and when a site changes we patch upstream while the schema we deliver to you stays fixed — so nothing downstream needs to be touched. If a change causes a genuine coverage gap we tell you which window was affected rather than quietly returning fewer records.
Yes, and we would rather you did. Every engagement starts with a free sample run against your own targets — your competitors, your catalogue, your keywords — and we hand back the actual dataset. Judging real output against your real requirements is the only useful evaluation, and it is a much better use of a first conversation than a slide deck.

Still have questions?

Talk to an engineer

Ready to Get Started?

Talk to us about your sources and volume. We'll return a sample dataset from your target sites before you commit to anything.

SOC 2 Type II
GDPR & CCPA compliant
99.9% uptime SLA

© 2026 PandorLabs, Inc. All rights reserved.