How ActPol measures competency, and what we publish
Every ActPol cycle asks the same two questions in the same order: what do you know, then what do you think. This page is about the first half — how a competency question gets written, checked, and scored — and what we commit to publishing about it, and when.
The rule for a competency question
A competency item must have an answer that a subject-matter expert who disagrees with us politically would still call correct. If a reasonable expert on the other side of the argument would dispute the answer, the question isn't a competency item — it belongs in the opinion stage, or it doesn't ship.
That's a bar, not a formality. “Is nuclear power safe?” fails it, because “safe” is a judgment. “Roughly what share of Hungary's electricity did Paks I generate last year?” passes it, because it has one defensible, cited answer.
Two questions, always in that order
Every respondent answers a short competency check on a topic before giving an opinion on it. The opinion is recorded either way — the competency check segments the results afterward, it never decides whose opinion counts.
Every question type — labelling a diagram, ordering steps, a multiple-choice fact, or estimating a number — includes an explicit “I don't know”, which scores zero but isn't treated as a wrong guess. Forcing a guess would measure willingness to guess, not knowledge, and that isn't something worth publishing.
How a score becomes a band
Each submission's competency score is a plain percentage — points earned over points available, for that topic — stored at submission time together with the version of the band thresholds used to compute it. That way, a later change to the thresholds can't quietly rewrite an old result out from under whoever already cited it.
The current thresholds — novice 0–39, familiar 40–64, informed 65–84, expert 85–100 — are a starting estimate, not a finding. Once a pilot cycle has run, we'll set them from the actual distribution of scores and freeze that version. We won't retune them to make a chart look better; the version number stamped on every result is what makes that checkable.
One place the scoring is genuinely counter-intuitive
Some questions ask you to put things in order — the stages of a process, the steps in a chain of cause and effect. Those aren't marked right-or-wrong. They're scored on how close your sequence is to the correct one overall, using a standard rank-correlation measure, so getting most of a sequence roughly right earns most of the marks.
That measure has a property worth stating before anyone discovers it and assumes we were hiding it: it scores the *whole* arrangement, so where your correct knowledge sits inside the list changes what it is worth. On a five-stage question, placing the last three stages in reverse can score higher than placing the first three correctly — because reversing a run still leaves most of the other pairs in the right relative order, while leaving two stages unplaced leaves every pair involving them wrong.
That's a real property of rank correlation, not a bug in our code, and it's why an ordering question will not let you submit a partial arrangement: a half-finished sequence isn't a half-answer, it's a differently-shaped one, and scoring it as though it were partial credit would be quietly unfair in a direction nobody could see. Place all of them, or answer "I don't know", which scores zero honestly and is never penalised.
It also has a consequence for how we write these questions, which we'd rather say than have inferred: because the arrangement is built from the front, a question whose middle is easier to know than its beginning will understate the people answering it. Ordering questions are written with an unambiguous first step for that reason.
Who writes a question, and who checks it
An item is drafted against a specific, recorded source — not a source found afterward to justify an answer already chosen. Whoever approves an item cannot be the person who wrote it; that's enforced directly, not just a policy someone could forget under deadline pressure.
Before an item is used, someone else has to be willing to accept the answer key — ideally an outside reviewer, and specifically someone who wouldn't be expected to agree with wherever the team is assumed to sit politically. We haven't run that outside review on any item yet; the first topic is waiting on it before it goes live, and no competency item ships without it.
What becomes public, and when
While a cycle is open, no answer key is readable outside the database that runs the survey — an open key would just be an answer sheet. The moment a cycle closes, that changes: the key and its citation become part of the public record for that item, for anyone to check.
That publication mechanism already exists in the database. What doesn't exist yet is a page to browse it from — every closed cycle's questions, answers and sources will get one.
What we won't say
Everyone who takes part is self-selected — nobody was sampled to represent a country, and adjusting the numbers afterward doesn't change that. So we won't publish a chart that reads “68% of Hungarians” or claims to speak for a country's population, ever — only what ActPol respondents, in that cycle, said.
What a cycle's numbers can support, and what we will publish, is how an answer differed between people who could back it up and people who couldn't — informed respondents differed from novice respondents by this many points on this question. That comparison holds up even though the sample isn't representative; a raw national-sounding figure wouldn't.
Common objections
- Isn't this just deciding whose opinion counts?
- No — every respondent's opinion is recorded and published, in total as well as by band. The competency check is an extra way of slicing the same results, not a filter on who gets counted.
- Doesn't your competency test just encode your own politics?
- That's the real risk, and it's why every item needs a citation, why the person who approves an item can't be its author, and why every answer key becomes public once its cycle closes. None of that is a guarantee — it's evidence anyone can go check, within weeks, not years.
- Competency probably correlates with education, income and age — isn't this just weighting toward people who are already better off?
- Probably true in part, and we'd rather publish it than argue around it. Every band's rough demographic makeup is published alongside its opinion. If the “expert” band on a topic turns out to be the same kind of person every time, that's a real finding about the instrument, and it gets reported, not smoothed away.
- Won't people just look up the answers?
- Some will — and a respondent who looks something up mid-question has, in a real sense, become a little more informed about it. What matters is that an open answer key can't be read straight off the page, and that how long someone took is recorded, so an implausibly fast run is something we can find and look at afterward.
What this doesn't try to do
No adaptive difficulty, no per-respondent ability score that follows someone across topics. Both are real techniques, and both would make the scoring harder to explain in one sentence — and being explainable is worth more here than being marginally more precise.