- A person parameter on the latent scale in the model.
- Rasch (1PL) constrains discrimination; 2PL lets items differ; 3PL adds guessing.
- They spend items near your level.
- Item exposure and coaching are security issues for high-stakes tests.
- A modelling assumption that items are not secretly one item in disguise.
- Inspection-time lite and odd-even switch lite are different item types.
What is theta?
A person parameter on the latent scale in the model. It is not a street IQ and not your 2-back hit count. 練習であり臨床IQではない。
Rasch versus 2PL?
Rasch (1PL) constrains discrimination; 2PL lets items differ; 3PL adds guessing. Manuals name the model. This site does not run one. 練習であり臨床IQではない。
Why adaptive tests feel shorter?
They spend items near your level. Short web quizzes that feel hard are not automatically IRT-adaptive. 練習であり臨床IQではない。
Can I “beat” IRT by practising items?
Item exposure and coaching are security issues for high-stakes tests. Practising this site’s toys is fine; it still is not a licensed theta. 練習であり臨床IQではない。
Local independence?
A modelling assumption that items are not secretly one item in disguise. Speeded 16-trial clouds violate lots of assumptions. That is expected for a toy. 練習であり臨床IQではない。
What to try?
Inspection-time lite and odd-even switch lite are different item types. Neither calibrates a published IRT bank. 練習であり臨床IQではない。
❓ よく一緒に検索される質問
What is construct validity?
Construct validity is the argument that a score reflects a named idea — working memory, not “being clever in general.” It is built from theory, patterns of correlations, and failed rival explanations. A 20-trial number-Stroop lite can feel like control. Feeling is not a validity coefficient. This site never reports a licensed IQ. 練習であり臨床IQではない。
What is construct validity? →What is criterion validity?
Criterion validity asks whether a score lines up with something you care about that is not the test itself — a later grade, a supervisor rating, or a Wechsler index. Concurrent means measured at the same time; predictive means later. A visual 2-back lite has no published criterion table. It is not an employment screen and not an IQ. 練習であり臨床IQではない。
What is criterion validity? →What is content validity?
Content validity is whether the item set covers the skill or knowledge you claimed. A spelling test with only the letter Q is a coverage failure. Shape-span lite samples a tiny visuospatial sequence. It does not cover “all spatial intelligence.” Experts, blueprints and reviews build content arguments. A weekend hackathon does not. 練習であり臨床IQではない。
What is content validity? →What is a confidence interval?
A 95% confidence interval is a procedure: if you repeated the study under the same model, 95% of such intervals would cover the parameter. It is not the probability that this one interval magically contains the truth after you saw it. Licensed IQ manuals turn a standard error into a range around a scaled score. Digit-span lite does not. Treat any integer from this site as a toy count. 練習であり臨床IQではない。
What is a confidence interval? →What is a g-loading?
In a factor analysis of several tests, a general factor (g) can appear. A subtest’s loading says how much it shares that factor in that sample and battery. Matrices often load high; some speeded motor tasks load lower. A loading is not your worth. N-back lite is not a published g-loading table. This site does not output g. 練習であり臨床IQではない。
What is a g-loading? →