ΨEmpiralis

Scientific Background

Every test on Empiralis is based on freely available, peer-reviewed questionnaires. For each test we document the original source, its license status, and what studies say about its strengths and limits.

What does "science-based" actually mean?

Psychological questionnaires are developed and evaluated against established psychometric criteria: reliability (measurement precision), validity (does the test actually measure what it claims to?), and norming (comparability against a reference sample). An online self-test can only partially meet these criteria, since e.g. the conditions under which it's taken aren't standardized. Our tests are meant for self-reflection, not diagnosis.

Advertisement

Scientific background by test

PHQ-9

The PHQ-9 was developed by Spitzer, Kroenke, and Williams (1999) based on the DSM-IV criteria for major depression and is one of the most extensively studied screening instruments in clinical psychology and primary care. Numerous studies show good internal consistency (Cronbach's alpha usually above .85) as well as good sensitivity and specificity for detecting depressive disorders. The PHQ-9 is a screening and monitoring tool, not a diagnosis by itself — that always requires a clinical conversation.

GAD-7

The GAD-7 was developed by Spitzer, Kroenke, Williams, and Löwe (2006) and is — alongside the PHQ-9 — one of the most extensively studied and widely used screening instruments in clinical psychology. Studies show good internal consistency (Cronbach's alpha usually above .89) and good sensitivity and specificity, not only for generalized anxiety disorder but also as a general indicator for panic disorder, social phobia, and PTSD. The GAD-7 is a screening and monitoring tool, not a diagnosis by itself.

WHO-5

The WHO-5 was introduced in 1998 by the WHO Regional Office for Europe and, with over a thousand validation studies in more than 30 languages, is one of the best-studied short instruments for psychological well-being. It's used, among other things, in primary care as the first step of a two-stage depression screening: scores of 50% or below are considered a sign of reduced well-being and, per common clinical practice, warrant further evaluation for depressive symptoms. Since 2024, the WHO itself holds the copyright to the WHO-5 and explicitly makes it available worldwide free of charge.

Burnout (CBI)

The CBI was developed in 2005 by Kristensen, Borritz, Villadsen, and Christensen as a freely available alternative to the commercial Maslach Burnout Inventory. It shows good internal consistency (Cronbach's alpha between .85 and .91) and correlates strongly with established burnout and exhaustion measures. The authors emphasize that burnout is not a stable personality trait but can change meaningfully over time -- retaking the test periodically can help make those changes visible. The Personal Burnout subscale (items 1-6) is also discussed in the research literature as a possible screening signal for autistic burnout. The authors themselves published normative data in 2004 (Borritz & Kristensen, nfa.dk): the only clinical cutoff they publish is a score of 50 or more as "a high degree of burnout" (in the PUMA project's comparison sample of roughly 1,900 employees in social/care professions, 16-22% reached this level depending on the subscale). The lower bound of the "average" range was additionally derived from the means/standard deviations reported there: Personal Burnout from a representative Danish population sample (N=1,498, M=32.7, SD=15.7), Work- and Client-related Burnout from the PUMA sample (Work: M=33.0, SD=17.7, N=1,910; Client: M=30.9, SD=17.6, N=1,752), since no general-population data exists for those two work-specific subscales.

Self-Compassion

The Self-Compassion Scale was developed in 2003 by Kristin Neff (University of Texas at Austin) and is the most widely used instrument in self-compassion research worldwide. Self-compassion comprises three complementary pairs of opposites: self-kindness versus self-judgment, experiencing common humanity versus isolation, and mindfulness versus over-identification with negative feelings. Studies consistently show good internal consistency across the six subscales, along with robust associations with lower depression and anxiety and higher well-being. The bands used here were derived from the means/standard deviations reported in the original study's Table 4 (Neff, 2003, N=391 U.S. college students): Self-Kindness M=3.05/SD=0.75, Self-Judgment M=3.14/SD=0.79, Common Humanity M=2.99/SD=0.79, Isolation M=3.01/SD=0.92, Mindfulness M=3.39/SD=0.76, Over-Identification M=3.05/SD=0.96 (each on the 5-point item scale). The "average" range corresponds roughly to one standard deviation around that mean for each subscale. Kristin Neff's own materials explicitly note that there are no official clinical norms for the SCS -- these bands are a reference point, not a diagnostic scheme.

Life Satisfaction

The Satisfaction With Life Scale was developed in 1985 by Ed Diener and colleagues as part of research into subjective well-being. It's considered the standard instrument for measuring global, cognitively-evaluated life satisfaction, and is deliberately distinguished from the emotional component of well-being (positive/negative affect). The scale shows high internal consistency across numerous studies and stable test-retest reliability appropriate for a trait measure, while still being sensitive to change -- for example, over the course of therapeutic interventions. The bands below follow the interpretation ranges Diener himself publishes ("Understanding SWLS Scores"), based on the scale's 5-35 point range.

Family Function

The Family APGAR was developed in 1978 by family physician Gabriel Smilkstein to quickly assess family function in a primary care setting -- named after the well-known APGAR score from obstetrics. The name is an acronym for the five areas it covers: Adaptation (support), Partnership (communication), Growth, Affection, and Resolve (shared time/connectedness). Studies consistently show high internal consistency (Cronbach's alpha .80-.90) and good test-retest reliability. Factor analyses suggest the five items mostly capture a single, global dimension of "satisfaction with family" rather than five separate constructs.

PSC-17 (Parent)

The PSC-17 is a short form developed by Gardner and colleagues (1999) from the original, 35-item Pediatric Symptom Checklist by Jellinek and Murphy (1988). In a comparison study against the much longer Child Behavior Checklist (CBCL) with 269 children, the PSC-17 showed comparable detection accuracy at a fraction of the length (Gardner et al., 2007). The three subscales correlate with different clinical presentations: attention with ADHD, internalizing symptoms with anxiety and depressive disorders, externalizing symptoms with oppositional and aggressive behavior. Like any brief screener, the PSC-17 does not replace a thorough evaluation -- its sensitivity is around 42-73% depending on the condition.

Parental Stress

The Parental Stress Scale was developed in 1995 by Berry and Jones as a shorter alternative to the 101-item Parenting Stress Index. Its distinctive feature: it deliberately captures not just demanding aspects, but also positive, enriching aspects of parenting -- both contribute to the overall picture of parental stress. Higher scores are associated in studies with lower parental sensitivity, more child behavior problems, and lower quality of the parent-child relationship -- the scale is used, among other things, to track change from parenting programs or family support. The bands used here follow a population-based Norwegian sample of 1,096 parents of one-year-olds (Nærde & Hukkelberg, 2020, PLOS ONE), which reports a mean of 31.0 (SD=7.27) on the 18-90 raw score range: the "average" range corresponds to roughly one standard deviation around that mean.

Vanderbilt (Parent, ADHD)

The Vanderbilt scales were developed by Mark Wolraich and colleagues based on DSM-IV criteria and validated in a large sample of referred children (Wolraich et al., 2003). They're among the most widely used ADHD screening instruments in U.S. pediatric primary care.

AUDIT

The AUDIT was developed by the World Health Organization as part of an international multi-center study (Saunders et al., 1993) and is today the most widely used screening instrument for hazardous drinking in primary care worldwide. It covers three areas: consumption (questions 1-3), dependence symptoms (questions 4-6), and alcohol-related problems (questions 7-10). Numerous international validation studies confirm good sensitivity and specificity for detecting hazardous use.

GAS-7 (Gaming)

The Game Addiction Scale was developed in 2009 by Lemmens, Valkenburg, and Peter for adolescents, based on classic addiction criteria (following Griffiths, among others): salience, tolerance, mood modification, relapse, withdrawal, conflict, and problems. The 7-item short form used here was validated in 2016 by Khazaal and colleagues in large French- and German-speaking samples from Switzerland and showed good internal consistency. Problematic gaming has been listed as "Gaming Disorder" in the WHO's ICD-11 since 2019 as its own diagnosis -- though that requires marked impairment over at least 12 months and a professional evaluation, which this brief screen cannot replace.

SCOFF

The SCOFF was developed in 1999 by Morgan, Reid, and Lacey in a British primary-care sample and published in the BMJ. In the original study, a threshold of two or more "yes" answers correctly identified roughly 85-100% of people with a diagnosed eating disorder (sensitivity), with similarly high specificity. SCOFF has since been translated into numerous languages and validated internationally. As a pure brief screen, it can produce false positives and false negatives and doesn't replace a thorough evaluation.

Attachment Style

The ECR-R was developed in 2000 by Fraley, Waller, and Brennan using item response theory, based on the original ECR by Brennan, Clark, and Shaver (1998), and is considered one of the psychometrically most robust instruments in attachment research. The combination of low attachment anxiety and low attachment avoidance is referred to in the research literature as a "secure attachment style"; high scores on one or both dimensions correspond to anxious-insecure, avoidant-insecure, or anxious-avoidant attachment patterns.

Loneliness

The UCLA Loneliness Scale was originally developed in 1978 by Russell, Peplau, and Ferguson, and revised by Daniel Russell into the third version used here in 1996. It's the most widely used research instrument in the world for measuring subjective loneliness and shows very good internal consistency across numerous studies (Cronbach's alpha usually above .90). Importantly: the scale measures the subjective feeling of loneliness, not the objective number of social contacts -- you can feel lonely in a crowd, and be alone without feeling lonely. The bands used here are based on the reference values reported in the original study (Russell, 1996, Table 2) for two adult samples using the full 20-item scale (college students, N=487, M=40.08, SD=9.50; nursing staff, N=305, M=40.14, SD=9.52): the "average" range corresponds to roughly one standard deviation around that pooled mean (M≈40.1, SD≈9.5). The original study does not define a formal clinical cutoff -- the scale is meant to be interpreted dimensionally, not categorically.

Big Five

This test is based on the International Personality Item Pool (IPIP), a freely available collection of personality items developed as a public alternative to commercial questionnaires like the NEO-PI-R. The 50 items used here correspond to the well-known IPIP representation of the five factors after Costa & McCrae (Goldberg, 1992) -- specifically the IPIP's 10-item scale for each factor. This is for self-reflection, not diagnostic classification. Studies on the Five-Factor Model show robust international replication of the factor structure, as well as good test-retest reliability over periods of several years.

Dark Tetrad (SD4)

Paulhus and Williams coined the term "Dark Triad" in 2002 for three socially unpopular but overlapping personality traits -- narcissism, Machiavellianism, and psychopathy -- which proved to be distinct but related constructs. Buckels, Jones, and Paulhus (2013) showed that everyday sadism (enjoyment of cruelty toward others) represents a fourth, empirically separable facet -- since then, researchers have spoken of the "Dark Tetrad." In 2020, Paulhus, Buckels, Trapnell, and Jones published the SD4 (Short Dark Tetrad) in the European Journal of Psychological Assessment as a compact, 28-item short form that captures all four traits with good internal consistency (Cronbach's alpha between .76 and .81), replacing the previous three-trait SD3 short scale from the same author team. Importantly: these traits are understood as normally distributed personality characteristics in the general population ("subclinical") -- they are conceptually distinct from the clinical diagnoses of Narcissistic or Antisocial Personality Disorder, or clinical psychopathy (which is assessed via, e.g., the clinician-administered PCL-R). The bands used here are based on the item means and standard deviations reported in the original study for a sample of 637 University of Winnipeg students, and serve only as a rough point of reference, not a clinical norm.

Self-Esteem

The scale was developed in 1965 by sociologist Morris Rosenberg and remains the most widely used instrument for measuring global self-esteem in psychological research today. Numerous studies across many countries and languages show very good internal consistency (Cronbach's alpha usually above .85) and stable test-retest reliability. The scale deliberately measures a one-dimensional, global sense of self-worth, not individual facets like a sense of competence or body image.

ASRS (Adult ADHD)

The ASRS was developed in 2005 by a WHO working group together with Ronald C. Kessler (Harvard Medical School) and colleagues, and is one of the most widely used screening instruments internationally for adult ADHD. The full version has 18 items; the six items used here ("Part A") were identified in the original study as the most statistically informative of all 18, and are validated as a standalone brief screener. Importantly: this questionnaire checks for current symptoms, but doesn't replace the evidence, required for an ADHD diagnosis, that similar difficulties were already present in childhood.

CAT-Q (Masking)

The CAT-Q was developed in 2019 by Laura Hull and colleagues at University College London, after clinical observations showed that many autistic people -- particularly women and adults diagnosed later in life -- camouflage their traits so successfully that classic screening instruments like the AQ don't reliably capture them. Camouflaging is associated with increased mental exhaustion, anxiety and depression, and delayed diagnosis. The questionnaire shows good internal consistency for all three subscales in validation studies and reliably distinguishes between autistic and non-autistic samples.

RAADS-R

The RAADS-R was published in 2011 by Ritvo and colleagues as a revised version of the original RAADS, and is one of the most widely used screening instruments in research and clinical practice for autistic traits in adults. It's organized into four domains: social relatedness, circumscribed interests, language, and sensory-motor. The original study proposed a total-score threshold of 65 points, above which autistic traits are likely (non-autistic control group mean: approximately 21 points). Importantly: both lower and higher scores occur in both autistic and non-autistic people -- in particular, pronounced "masking" (see the CAT-Q) can lower the score. This test can therefore never replace a diagnosis, only provide an indication that further professional evaluation may be worthwhile.