Not every skill belongs in an MCQ, and not every assessment method predicts job performance equally well. Decades of industrial-organisational psychology research have produced a clear hierarchy of assessment validity — the ability of a method to predict how someone will actually perform in a role. Choosing the wrong method for a skill wastes candidate time, recruiter time, and produces data you cannot trust for decisions. This guide covers the main skills assessment methods, their strengths and limits, and how to combine them for a practical, scalable assessment process.
A technically perfect question on a badly chosen assessment format produces misleading results. A software engineer who can explain every design pattern in an MCQ format might be a poor debugger under time pressure. A salesperson who scores poorly on a written aptitude test might be extraordinarily effective in person. Assessment validity — the degree to which a method actually predicts performance — varies dramatically between methods and skills.
Knowledge tests (including MCQ-based assessments) measure what someone knows: facts, concepts, procedures, and domain vocabulary. They are fast to administer, cheap to scale, and easy to score objectively. They are highly valid for roles where knowledge itself determines performance — medical, legal, compliance, technical domains where getting facts wrong has direct consequences.
Work samples ask candidates to do a representative piece of actual work. A writing sample for a content role, a data analysis task for an analyst, a code review exercise for a developer. Research consistently shows work samples have high predictive validity because the correlation between performing the task in assessment and performing it on the job is direct, not indirect. The tradeoff is cost: work samples take time to design, administer, and evaluate.
A structured interview asks every candidate the same questions, in the same order, with a pre-defined scoring rubric for each answer. Unstructured interviews — where interviewers improvise questions — have low predictive validity and high evaluator bias. Structured behavioral interviews ("Tell me about a time when...") and situational interviews ("What would you do if...") are both significantly more predictive of job performance.
General cognitive ability (GCA) tests — verbal reasoning, numerical reasoning, abstract reasoning — are among the strongest single predictors of job performance across a wide range of roles. They measure learning ability and problem-solving speed rather than current knowledge. The implication: a candidate who scores well on cognitive aptitude but has no current domain knowledge may outperform a domain expert who scores poorly, given enough time to learn.
Manager-rated performance reviews assess skills over time based on observed work. They are valuable for development decisions on existing employees and for identifying high performers. They are unreliable as a selection tool for new hires because there is no performance to observe. For existing employees, their validity improves significantly when the rating is structured (rubric-based, evidence-required) rather than free-form opinion.
Validity in assessment research is measured as a correlation between assessment score and subsequent job performance, expressed as a number between 0 and 1. A validity of 0.5 means the assessment explains 25% of the variance in job performance — meaningful, but not deterministic. For context, unstructured interviews have validity around 0.2; structured interviews around 0.4–0.5; cognitive ability tests around 0.5; work samples around 0.5; combinations of methods can reach 0.6+. No single method predicts performance perfectly.
The practical implication of these numbers is that any single assessment method leaves significant variance unexplained. The case for combining methods is not theoretical — it is empirical. Adding a structured interview to a cognitive ability test consistently produces better predictions than either method alone. This is why well-designed hiring processes use at least two complementary methods.
Online assessments shine in three scenarios: filtering large candidate pools where you need to eliminate clearly unfit applicants before spending interviewer time, verifying that training content was absorbed, and standardising baseline knowledge measurement across a diverse group. In all three cases, the knowledge test is screening for a minimum threshold, not making a final selection decision.
The strongest assessment processes combine at least two methods. For volume hiring, the typical stack is: cognitive aptitude screen online, then a domain knowledge check, then a structured interview for shortlisted candidates. For training, it is: pre-assessment, training delivery, post-assessment, then manager observation. Each layer answers a different question and reduces the risk of false positives from any single method.
Higher-validity methods (work samples, structured interviews) cost more per candidate to administer. This is the core trade-off in assessment design. The practical solution: use lower-cost, lower-fidelity methods early in the funnel to screen the majority, and reserve higher-cost, higher-fidelity methods for a smaller shortlist. Online aptitude and knowledge tests are specifically valuable in this role — they are cheap enough to run on every applicant and accurate enough to make meaningful cuts before you spend interview hours.