Profila Sign up

The 42-True corpus

What people asked for,
and whether we
got it right.

A declared-intent dataset where the person who wrote the request also judged the answer.

42-True Not yet licensable

Status

The corpus is live and growing. It is not yet licensable.

The 42-True corpus runs in production today. It is safe to publish — anonymity is enforced in code and independently re-checked — but it is not licensable, and the blocking criteria are named in full on the Readiness tab rather than buried in a schedule.

We are not publishing release figures yet. Below roughly a hundred contributors a corpus describes a handful of individuals rather than a population, so counts and rates from it would be precise and misleading at the same time. The measured figures — records, contributors, unmet-demand and correction rates, category coverage — are published once the corpus passes 100 contributors. Until then we will describe the mechanisms, and give you the current numbers directly if you ask.

Nothing here is an offer. The licensing models are proposals, no pricing is stated, and third-party model training is a permission that does not yet exist on any record. We would rather you evaluated this on evidence when the evidence means something than took a pitch on trust now.

42-True

What it is

A record of what people said they wanted, and whether they said it was right

Each record is a four-part chain.


          declaration  →  classification  →  match  →  outcome (+ the person's own verdict)
        
  • Declaration — a person, in their own words, stating an intent. “a family estate car under 40k, petrol, in Zurich.” Not inferred from browsing. Typed, deliberately, by someone who wanted to be shown something.
  • Classification — what our system understood that to mean, expressed as an IAB Content Taxonomy v3.1 code, with the model’s confidence and the exact dictionary version used.
  • Match — what was actually served in response, or the explicit fact that nothing could be.
  • Outcome — what the person then did, on a nine-rung behavioural ladder, plus their explicit verdict on whether the match was right.

That last element is the part that does not exist elsewhere at scale. Every record carries a human answer to “is this a good match?” from the person who wrote the declaration — not a proxy, not an inferred label, not a crowd-worker rating a screenshot.

Every record is verified by the declarant. Not a sample. All of them.

42-True

What it is

Four things a record can tell you that an ad log cannot

Stated intent, not inferredThe person typed what they wanted. No behavioural reconstruction, no lookalike modelling, no cookie graph.
A first-party verdictThe declarant judged the match. The label comes from the only person who actually knows.
Unmet demandDeclarations that could not be filled are recorded as first-class data. An ad network cannot see these — it only ever sees the impressions it sold.
CounterfactualsWhat the ranker considered and did not serve, with the reason each candidate was eliminated.
42-True

Why it differs

The intent-data market has a labelling problem

Most commercially available “intent data” is inferred. A person visits three pages about electric vehicles; a model concludes they are in-market for one. The label is a guess, and it is never checked against the person. Errors are invisible because nothing ever contradicts them.


          Inferred intent:   behaviour  →  model  →  label            (never checked)
42-True:           person     →  words  →  match  →  person's verdict
        

The consequences compound. A model trained on inferred intent learns to reproduce the inference engine that generated the labels, not the intent itself. Its errors are correlated with that engine’s errors, and no amount of scale fixes it — more data from the same generator gives you a more confident version of the same mistake.

The label and the intent come from the same human, minutes apart. When the match is wrong, we know — because they told us, and the record says so.

42-True

Why it differs

The disagreements are the point

Cross the person’s verdict against what they actually did, and the cells where the two contradict each other are populated — not theoretical.


                                 no_engagement   click   deep_engagement   dismiss
  yes_match                  ·            ·            ·
  no_match                                ·                          ·
  no_response  (nothing was shown, so nothing was judged)
        

The cell that matters is no_match | click. Those are people who clicked something and then said it was wrong. In any click-based dataset they are positive training examples. Here they are labelled as the errors they are.

The mirror cell is yes_match | no_engagement: content the declarant confirmed was right and then ignored. That is a creative problem, not a targeting one — and no dataset that infers relevance from behaviour can tell the two apart.

Those cells are the whole argument in miniature. Engagement is not relevance, and only the declarant can tell you the difference. How many records sit in each cell is a figure we publish once the corpus passes 100 contributors — below that it is a description of a few people, not a distribution.

42-True

Why it differs

Unmet demand — the structurally unavailable signal

When someone declares an intent and the search returns nothing usable, that is recorded as a record with match.kind: "none" and the reason no_ads_found.

A material share of every release is unmet demand — the rate is one of the figures we publish once the corpus passes 100 contributors, because a percentage drawn from a handful of people describes those people rather than a market.

No advertising network can produce this. A network observes the impressions it sold; it has no representation for “a person wanted this and nobody could supply it.” For a brand deciding what to stock, which markets to enter, or where a competitor is under-serving, that is often the most valuable row in the table.

42-True

Why it differs

Correction chains — the premium shape

When a person says a match is wrong and then re-states what they meant, we capture the whole trajectory.


          declaration A  →  match  →  "no, that's wrong"  →  declaration B (refined)  →  match  →  "yes"
        

This is correction data: the before, the after, the human judgement that connects them, and the taxonomy delta between the two. It is the highest-value shape in preference learning, because it contains the direction of the error and not merely its existence.

This is an early-stage corpus and the Readiness tab is explicit about that — but the mechanism is live and producing real chains today, not a planned feature. Chain counts and resolution rates join the published figures at 100 contributors.

42-True

How it is made

The lifecycle of one record

Profila is a consumer application built on a deliberate inversion of the advertising model: the person declares what they are looking for, and brands respond. The corpus is a by-product of that product working normally — nothing is collected for the corpus that the product does not already need.

  1. Declare. The user writes an intent in their own words and selects a market, a timeline, and what kind of result they want.
  2. Classify. An LLM maps the declaration to an IAB Content Taxonomy v3.1 code and reports a confidence. The dictionary version and its content hash are recorded.
  3. Search. An agent constructs a query and searches. Candidates pass four independent gates — a deterministic relevance prefilter, a hard country filter, an LLM relevance floor, and a post-fetch geographic confirmation — plus a host-diversity cap.
  4. Match. Surviving candidates become content in the user’s feed. If nothing survives, the attempt is recorded as unmet demand rather than discarded.
  5. Verdict. Each card carries “Is this a good match?” — a binary Yes/No. The user’s answer is the relevance label, and they can change it afterwards.
  6. Engage. Behaviour is recorded on the ladder: impression, click, deeper engagement, lead, and so on.
  7. Refine. If the user says “no”, they can re-state the intent. The new declaration is linked to the old one, forming a correction chain.
  8. Export. At release time the four-tuple is assembled, anonymised, validated against 14 rule families, signed, and written to the release.
42-True

How it is made

Product decisions made to keep the labels clean

These are worth stating because they cost us volume.

  • The verdict is binary. Yes or No. There is no “unsure” option — it invites overthinking and produces a bucket that means nothing consistent. The ability to change the answer afterwards handles genuine mistakes. No answer remains a distinct state, recorded as the absence of a verdict rather than as a third button.
  • There is no separate “dismiss” control. “No” already is the dismissal. A second negative path would compete with the primary one and split the signal.
  • Dwell time is bucketed, never raw. Millisecond dwell is a behavioural fingerprint. The corpus carries a coarse band — enough to distinguish “read it” from “scrolled past”, which is all the behavioural rung needs.
  • We never inject commercial bias into a query. What the user asked for is what gets searched. The only thing that steers toward commercial results is the user’s own explicit choice of result type.
42-True

The research behind it

The protocol this corpus is a by-product of

42-True is the protocol: people declare what they want, and the network resolves those declarations without observing the people who emit them. The open research — including the Large Meaning Model the corpus is meant to train — is on Profila Research Labs.

These pages are the other half: the dataset that protocol produces, described for someone deciding whether to license it.

42-True

Diligence

We would rather be checked than believed

A sample release is available under NDA, along with the signing public key, the validation rule set and the most recent re-validation report. The Readiness tab states what is blocking a licence today, and Licensing describes the shapes we think fit the data.