Appen is an AI training data company that sells human-produced and human-validated data, data annotation and model evaluation to teams building AI systems. It was founded in Sydney in 1996, initially focused on speech and language data for early NLP systems. Today it is listed on the Australian Securities Exchange under the code APX. Appen describes its job as delivering the expert-validated data that trains frontier models, so that AI systems understand nuance, context and complexity at scale.
The company works on two sides of the same market:
Appen reports more than 1 million vetted contributors, 170+ countries represented and 235+ supported languages. In 2018, the year of its ASX listing, Appen acquired Figure Eight to expand its platform capabilities.
Appen groups its custom data work into six product lines, each built for one part of modern model development:
For subject-matter-expert RLHF, Appen says preference rankings and comparative feedback come from verified PhDs, MDs and JDs in medicine, law, science and finance, arguing that standard crowdsourced annotation cannot judge whether a clinical diagnosis or a legal argument is correct. Its regulatory audits evaluate model outputs against the EU AI Act, the NIST AI RMF and organisational ethics frameworks.
A separate coding line offers repository-level datasets for training and evaluating LLMs and coding agents, with 60+ environments, 50+ languages and 500 verifiable RL tasks. The environments are described as private and kept out of public code indexes, so evaluations can run without known training overlap.
ADAP is Appen's data annotation platform, and Appen says it merges automation and human oversight to deliver data across a wide range of modalities and AI use cases. Its main capabilities, as listed by Appen:
Appen's catalog holds 596 datasets across eight categories, from RL tasks and verifiers to speech, code and enterprise data. Every dataset ships with a spec sheet documenting source, consent, annotation method, known limitations and license terms. Appen frames the choice this way: off-the-shelf sets cover common categories under standard conditions and deliver in days, while custom collection is used when a project needs specific demographics, controlled acoustic environments or specialist domain coverage.
Appen also runs a supply-side program in which companies license the workflows and operational know-how they already run on to AI labs building new models, while Appen handles extraction, anonymisation and the conversations with the labs. It scopes three categories of operational data: operational documentation, decision and escalation records, and systems of record.
Appen's published case studies show the range of work it takes on; the figures below are the company's own.
Appen fits:
It is a weaker fit for:
Mobile tasks center on taking pictures of objects, recording video or recording audio, and Appen recommends a computer for some registration and project selection steps before work begins in the app.
Appen pricing is not published for buyers. In the official pages checked for this review, including the full sitemap, there is no price list, plan table or self-serve checkout, and the AI data team's contact form is marked for business and sales inquiries only.
The standard catalog license is perpetual and non-exclusive with commercial training rights, and the license type is listed on each dataset's spec sheet. For enterprise data, each partnership is quoted individually, terms are agreed in writing before anything is extracted, and payment is made on acceptance. Appen covers extraction, anonymisation, preparation and transfer, with no upfront costs, no setup fees and no deductions from what partners are paid. Value rises with how rare the operational knowledge is, how connected the source systems are, and how much data there is and how far back it goes.
CrowdGen advertises $100+/hr as top hourly pay, with a footnote that the pay rate is based on location and skills. Medical clinicians and legal experts get the highest domain caps, up to $450/hr and up to $400/hr, with software engineers at up to $150/hr and linguists at up to $95/hr. Finance caps run from up to $150/hr for investment managers to up to $180/hr for FP&A analysts, while healthcare administrators top out at up to $90/hr. These are ceilings, not typical earnings. CrowdGen says rates are stated clearly before work begins. Appen also says it will never ask contributors to pay anything to join its projects.
Services that overlap with Appen, each summarized by its own positioning:
Appen's own mix adds a licensable dataset catalog, a program for licensing enterprise operational data, and physical-AI data captured, synchronized and annotated end-to-end in purpose-built facilities.
Yes. CrowdGen describes itself as the contributor platform of Appen, and Appen's scam-awareness page lists crowdgen.com and app.crowdgen.com among its official websites.
Appen says it will never ask people to send money to apply for or start a job, never sends checks in advance, and sends official emails only from @appen.com addresses.
Most earnings are stated in US dollars and paid in local currency, with withdrawals through bank transfer, PayPal, Payoneer, Airtm and other methods.