All questions
Question 1
A test is redesigned so the new norm group yields mean 100 and SD 15. What process is this?
- Standardization (renorming), because scores are recalibrated to a representative sample so the distribution has mean 100 and SD 15. (correct answer)
- Test-retest reliability, because changing the norm group guarantees students will obtain the same score every time they retake the test.
- Construct validity, because setting mean 100 and SD 15 proves the test measures intelligence rather than motivation or schooling.
- Fixed-score calibration, because IQ is unchangeable and renorming is unnecessary once the first norms are established for a population.
Explanation: Standardization or renorming is the process of establishing new norms for a test by administering it to a large, representative sample and adjusting scores to fit a predetermined distribution. For IQ tests, this typically means calibrating scores so the population mean equals 100 and standard deviation equals 15. This process ensures test scores remain interpretable and comparable across time and populations. Renorming is necessary periodically because population performance can shift (as seen in the Flynn effect). The standardization sample should represent the population for whom the test will be used, including appropriate demographic diversity. Without proper standardization, raw scores would be meaningless - it's the comparison to the norm group that gives IQ scores their interpretive value.
Question 2
A school worries an IQ test underestimates students who are bilingual; what concept best addresses this concern?
- Cultural and linguistic bias, because language demands can depress scores unrelated to reasoning, lowering validity for some groups. (correct answer)
- Higher reliability, because bilingual students' lower scores prove the test is consistent and therefore measures intelligence accurately.
- Spearman's g, because bilingualism interferes with the single intelligence factor, which is otherwise fixed and uniform across tasks.
- Fixed IQ, because bilingual students should score the same regardless of language exposure if intelligence is unchangeable.
Explanation: Cultural and linguistic bias represents a significant validity threat when tests place heavy language demands on examinees whose primary language differs from the test language. Such tests may underestimate reasoning abilities because poor performance could reflect language barriers rather than cognitive limitations. This introduces construct-irrelevant variance that threatens valid interpretation of scores. Bilingual students might understand the underlying concepts being tested but struggle with linguistic presentation, leading to scores that don't accurately reflect their cognitive abilities. Addressing this concern requires careful test development, including item review for cultural content, consideration of alternative assessment formats, and sometimes separate norms for different linguistic groups. The Flynn effect demonstrates that environmental and cultural factors significantly influence test performance. Understanding cultural bias is essential for fair testing practices and appropriate score interpretation, particularly in diverse educational settings where students bring varied linguistic and cultural backgrounds.
Question 3
A student generates many novel uses for a brick on a test; which Sternberg component is most involved?
- Creative intelligence, because producing novel, useful ideas and approaching problems in original ways is central to this component. (correct answer)
- Analytic intelligence, because generating unusual ideas mainly reflects evaluating arguments and applying formal logic to one correct solution.
- Validity, because many ideas show consistent scoring across time, which is the definition of validity rather than reliability.
- Fixed IQ, because creativity cannot be trained and therefore reflects an unchangeable intelligence level set at birth.
Explanation: According to Sternberg's triarchic theory, creative intelligence involves the ability to generate novel, useful ideas and approach problems in original ways. Tasks requiring students to think of unusual uses for common objects (like a brick) specifically tap into creative thinking abilities by demanding original, innovative responses rather than conventional applications. This component of intelligence is distinct from analytical intelligence (academic problem-solving) and practical intelligence (real-world application). Creative intelligence is particularly important for innovation, artistic endeavors, and adapting to novel situations. The Flynn effect suggests that environmental factors can influence cognitive development, and research indicates that creative abilities can be developed through appropriate instruction and experiences. Understanding creative intelligence helps explain why some individuals excel at generating original ideas even if they don't perform as well on traditional academic measures. This perspective has influenced educational practices to include more opportunities for creative expression and divergent thinking.
Question 4
A researcher correlates an IQ test with later job performance ratings; which type of validity is being assessed?
- Test-retest reliability, because correlating scores with later ratings checks whether the test gives identical results across time.
- Construct validity, because any correlation with any outcome automatically proves the test measures intelligence as a psychological construct.
- Criterion-related validity, because the test's usefulness is evaluated by how well scores predict an external criterion like job performance. (correct answer)
- Fixed IQ, because predicting job performance shows intelligence is unchangeable and therefore determines workplace success completely.
Explanation: Criterion-related validity (also called predictive validity when future outcomes are involved) assesses whether test scores correlate with external criteria that the test should theoretically predict. By correlating IQ scores with later job performance ratings, the researcher is evaluating whether the test successfully predicts a real-world outcome it claims to be relevant for. This differs from test-retest reliability (consistency over time) or construct validity (whether the test measures the theoretical construct). Strong criterion-related validity would show significant positive correlations between IQ scores and job performance, supporting the test's practical utility. This type of validity is crucial for justifying test use in selection contexts.
Question 5
Average IQ scores rise across generations while tests remain standardized to mean 100, SD 15. What is this called?
- Stereotype threat, because anxiety from group stereotypes causes long-term generational increases in average IQ performance.
- The Flynn effect, because population-level IQ performance tends to increase over time across successive cohorts. (correct answer)
- Test-retest reliability, because repeating the same IQ test over decades yields higher scores due to consistency.
- Fixed intelligence, because rising averages show IQ is genetically determined and cannot be changed by environment.
Explanation: This phenomenon is known as the Flynn effect, named after researcher James Flynn who documented consistent increases in IQ scores across generations in many countries. Despite tests being continually re-standardized to maintain a mean of 100 and standard deviation of 15, raw scores have increased approximately 3 points per decade. This means that someone scoring 100 today would likely score higher on an IQ test from decades ago. The Flynn effect suggests environmental factors play a significant role in intelligence test performance, possibly including improved nutrition, education, test familiarity, and cognitive complexity in modern life. This finding challenges notions of fixed intelligence and highlights the importance of considering cohort effects when interpreting IQ scores. The effect has important implications for how we understand intelligence and its measurement across time.
Question 6
A 10-year-old has a mental age of 12. What is the ratio IQ?
- 83, below the mean
- 100, at the mean
- 120, above the mean (correct answer)
- 130, likely gifted
Explanation: Divide mental age by chronological age and multiply by 100: 12 / 10 = 1.2, times 100 gives 120, which is above the mean of 100. The tempting wrong answer 83 comes from reversing the fraction (10 / 12), but ratio IQ always puts mental age over actual age, not the other way.
Question 7
Which finding would most weaken the claim that intelligence is a single general ability?
- Scores correlate across tests
- Heritability of IQ is high
- Fluid and crystallized co-vary
- High spatial, low verbal skill (correct answer)
Explanation: Seeing high spatial but low verbal skill in the same person shows abilities can separate, undercutting one general factor. If all intelligence were one g, performance would tend to be similar across domains. Tempting? Fluid and crystallized co-varying actually supports the claim, so it wouldn't weaken it.
Question 8
Which use best exemplifies an aptitude test?
- Predicting new skill mastery (correct answer)
- Certifying course mastery
- Ranking prior academic grades
- Diagnosing a reading problem
Explanation: Aptitude tests measure your potential to acquire a new skill, so predicting new skill mastery is the clearest fit. The tempting wrong answer is diagnosing a reading problem, because that is a diagnostic test of current ability, not a forecast of future performance. Certifying mastery and ranking grades also look backward at what you have already learned.
Question 9
In this population, the heritability of IQ is .50. Which statement follows?
- A person's IQ is half genetic
- Group IQ gaps are genetic
- Half of IQ variance is genetic (correct answer)
- Genes determine half of IQ
Explanation: Heritability of .50 means that, in this population, genetic differences account for half of the variation in IQ scores. It is a population-level statistic about variance, not about any one person. The tempting error is thinking a person's IQ is half genetic, but heritability says nothing about an individual's IQ makeup.
Question 10
A new aptitude test is reliable but does not predict achievement. Which validity is low?
- Content validity
- Predictive validity (correct answer)
- Construct validity
- Face validity
Explanation: Because the test fails to forecast later achievement, the validity that is low is predictive validity, which measures how well a test score predicts a future criterion. The tempting wrong answer is content validity, but that is about whether the test's items cover the relevant material, not about predicting outcomes.
Question 11
An intelligence test includes puzzles, vocabulary, and spatial tasks; what is the most defensible claim about what it measures?
- It samples certain cognitive skills and can be useful for prediction, but it does not capture every aspect of intelligence or potential. (correct answer)
- It measures all forms of intelligence completely, because standardization to mean 100, SD 15 guarantees comprehensive assessment.
- It is valid if it is reliable, because consistency across administrations is the same thing as measuring the intended construct.
- It proves intelligence is fixed, because performance across different item types reveals unchangeable cognitive capacity.
Explanation: An intelligence test including diverse cognitive tasks (puzzles, vocabulary, spatial tasks) samples certain cognitive skills and can provide useful information for prediction and description, but it cannot capture every aspect of human intelligence or potential. This represents a balanced, defensible view that acknowledges both the utility and limitations of IQ testing. Such tests may reflect Spearman's g factor and show predictive validity for academic and occupational outcomes, but they don't measure all forms of intelligence proposed by theorists like Gardner or Sternberg. The Flynn effect demonstrates that test performance can change over time due to environmental factors. Modern understanding recognizes that intelligence is multifaceted and that single tests, regardless of their breadth, provide limited perspectives on human cognitive abilities. This view supports using intelligence tests as one source of information among many, rather than as definitive measures of fixed intellectual capacity. Comprehensive assessment often requires multiple measures and consideration of diverse abilities and contexts.
Question 12
A test has high reliability but low validity; which outcome is most plausible?
- It produces consistent scores but measures the wrong construct, such as reading speed instead of reasoning ability for its intended use. (correct answer)
- It measures exactly what it claims, but scores vary widely each time because validity causes inconsistency across administrations.
- It supports Gardner's theory, because multiple intelligences guarantee a test can be consistent yet still not measure intelligence.
- It proves IQ is fixed, because a reliable test cannot be invalid if intelligence is stable and unchangeable.
Explanation: A test with high reliability but low validity consistently measures something, but not what it's intended to measure. For example, a test designed to measure reasoning ability might consistently measure reading speed instead due to heavy text demands. The scores would be stable across administrations (reliable) but wouldn't reflect the intended construct (invalid). This situation demonstrates why reliability and validity are distinct psychometric properties. High reliability is necessary but not sufficient for validity - consistency doesn't guarantee accuracy. Understanding this distinction is crucial for test development and interpretation. The Flynn effect shows that even valid measures may need periodic renorming, and different theories of intelligence (Spearman's g, Gardner's multiple intelligences, Sternberg's triarchic) suggest various approaches to valid measurement. This scenario highlights the importance of construct validation and careful consideration of what psychological tests actually measure versus what they claim to measure.
Question 13
A psychologist claims one general factor influences performance across many mental tasks. Which intelligence theory is this?
- Gardner's multiple intelligences, because independent modules like musical and bodily-kinesthetic abilities are separate and not strongly correlated.
- Sternberg's triarchic theory, because intelligence is best explained by analytical, creative, and practical components working in different settings.
- Spearman's g, because a single general intelligence factor helps explain positive correlations among diverse cognitive test performances. (correct answer)
- Fixed-IQ trait theory, because a single score permanently determines all cognitive abilities and cannot be influenced by environment or learning.
Explanation: Spearman's g (general intelligence) theory proposes that a single underlying factor influences performance across diverse cognitive tasks. Charles Spearman observed that people who perform well on one type of mental test tend to perform well on others, suggesting a common factor. This g factor represents general cognitive ability that contributes to all intellectual tasks, though specific abilities (s factors) also exist. The theory explains why cognitive test scores tend to correlate positively - they all tap into this general intelligence to some degree. This contrasts with theories proposing multiple independent intelligences (like Gardner's) or those emphasizing different types of intelligence (like Sternberg's triarchic theory). The g factor remains influential in intelligence research and psychometric testing.
Question 14
A student excels at composing music but is average on math and vocabulary tests. Which theory best fits this pattern?
- Spearman's g, because a single general factor should produce uniformly high performance across academic tasks and artistic performance together.
- Reliability theory, because consistent scoring across graders proves the student's musical talent is a valid measure of overall intelligence.
- Fixed-IQ view, because an IQ score permanently determines musical ability, so high music skill must imply high scores in all subjects.
- Gardner's multiple intelligences, because strengths can appear in specific domains like musical intelligence without equally strong academic abilities. (correct answer)
Explanation: Gardner's theory of multiple intelligences proposes that intelligence consists of several independent abilities or "intelligences" that operate separately. This theory explains why someone can excel in one domain (like musical intelligence) while showing average performance in others (linguistic or logical-mathematical). Gardner identified eight intelligences including musical, bodily-kinesthetic, interpersonal, and intrapersonal, arguing that traditional IQ tests only measure a narrow range of abilities. The theory challenges the notion of a single g factor determining all cognitive performance. It has been influential in education, encouraging recognition of diverse talents, though it faces criticism for lack of empirical support and difficulty in measurement. The student's profile of exceptional musical ability with average academic performance exemplifies Gardner's concept of domain-specific intelligences.
Question 15
An IQ test is standardized to mean 100, SD 15; what score is two SDs above average?
- A score of 130, because it is two standard deviations above 100 using an SD of 15 on the typical IQ scale. (correct answer)
- A score of 115, because one standard deviation above the mean indicates very superior intelligence and strong test reliability.
- A score of 145, because Gardner's multiple intelligences predicts many people exceed 140 when schools teach to their strengths.
- A score of 160, because IQ is fixed and extreme scores reflect unchangeable genetic endowment rather than measurement conventions.
Explanation: On a standardized IQ scale with mean 100 and standard deviation 15, calculating scores at specific standard deviation distances is straightforward arithmetic. Two standard deviations above the mean equals 100 + (2 × 15) = 130. This demonstrates how IQ scores are distributed on the normal curve, where approximately 95% of scores fall within two standard deviations of the mean. Understanding this standardization is crucial for interpreting IQ scores in both clinical and educational settings. The Flynn effect shows that population averages can shift over time, requiring periodic renorming to maintain the mean at 100. Gardner's multiple intelligences theory and concepts about fixed intelligence don't change the mathematical relationship between standard deviations and score interpretation.
Question 16
A student excels at music and interpersonal skills but average on logic puzzles; which theory best fits this profile?
- Gardner's multiple intelligences, emphasizing distinct abilities like musical and interpersonal intelligence that may not align with logical-mathematical performance. (correct answer)
- Spearman's g, because all cognitive abilities are driven by one factor, so strengths must appear equally across domains.
- High reliability, because consistent performance in music indicates the test is valid and measures intelligence rather than practice effects.
- Fixed IQ doctrine, because domain strengths show intelligence is permanent and cannot be influenced by instruction or cultural opportunities.
Explanation: Gardner's theory of multiple intelligences proposes that intelligence consists of several relatively independent abilities, including musical, interpersonal, logical-mathematical, linguistic, and others. A student who excels in music and interpersonal skills but performs averagely on logic puzzles exemplifies this theory's core premise that individuals can have distinct strength profiles across different intellectual domains. This contrasts with Spearman's g theory, which emphasizes a single general intelligence factor underlying all cognitive abilities. The Flynn effect describes population-level changes over time, and modern research shows that intelligence can be influenced by education and environment. Gardner's theory has been influential in education, encouraging recognition of diverse talents and alternative approaches to instruction that capitalize on different intellectual strengths.
Question 17
A test predicts first-year college GPA from high school juniors' scores; what validity is being evaluated?
- Predictive validity, because the test is judged by how well earlier scores forecast a later real-world outcome like college GPA. (correct answer)
- Internal consistency reliability, because predicting GPA requires the test to measure the same trait at two different time points.
- Gardner's theory, because GPA prediction depends on multiple intelligences, each standardized to mean 100, SD 15 independently.
- Fixed IQ certainty, because accurate prediction proves intelligence cannot change and therefore future academic performance is predetermined.
Explanation: Predictive validity is demonstrated when test scores successfully forecast future outcomes or performance in relevant real-world situations. A test that predicts first-year college GPA from high school scores shows it can anticipate future academic success, which is a key form of criterion-related validity. This type of validity is particularly important for selection and placement decisions in education and employment. The time gap between test administration and outcome measurement is what distinguishes predictive validity from concurrent validity, where measures are taken simultaneously. Strong predictive validity provides confidence that test scores have practical utility beyond the testing situation itself. This concept is fundamental to understanding why aptitude tests are valuable - their worth lies primarily in their ability to forecast future performance rather than just describe current abilities.
Question 18
Which statement best distinguishes reliability from validity in psychological testing?
- Reliability is consistency of measurement, while validity is whether the test measures what it claims to measure for the intended purpose. (correct answer)
- Reliability means the test measures the intended trait, while validity means scores are stable across time and across forms.
- Reliability is explained by Gardner's theory, while validity is explained by Spearman's g factor across all cognitive domains.
- Reliability and validity both prove IQ is fixed, since a stable mean 100, SD 15 scale cannot change with environment.
Explanation: Reliability refers to the consistency or stability of test scores - whether a test produces similar results when administered repeatedly under similar conditions. Validity, on the other hand, concerns whether a test actually measures what it claims or purports to measure for its intended use. A test can be highly reliable (consistent) but invalid (measuring the wrong thing), but a test cannot be valid without being reliable to some degree. These are distinct but related psychometric properties. Understanding this distinction is crucial for test interpretation and development. For example, a scale that consistently reads 5 pounds heavy is reliable but not valid for measuring true weight. The Flynn effect demonstrates that even reliable tests may need periodic renorming, and various intelligence theories (Spearman's g, Gardner's multiple intelligences, Sternberg's triarchic) describe different conceptualizations of what intelligence tests might validly measure.
Question 19
Scores rise over decades on the same IQ test norms; what concept describes this population-level increase?
- The Flynn effect, describing generational increases in average test performance that can shift norms despite stable mean 100, SD 15 scaling. (correct answer)
- Stereotype threat, because anxiety about confirming stereotypes makes later generations score higher when pressure is removed.
- Reliability, because rising scores across decades prove the test is consistent and therefore measures intelligence validly.
- Fixed IQ inheritance, because rising scores show genetic intelligence is steadily increasing and unaffected by environment or education.
Explanation: The Flynn effect describes the well-documented phenomenon of rising average IQ scores over several decades within populations. This generational increase in test performance has been observed across many countries and requires periodic renorming of tests to maintain the standard mean of 100. The Flynn effect demonstrates that population-level cognitive performance can change over time, likely due to factors such as improved education, nutrition, healthcare, and environmental complexity. This finding challenges simplistic views of fixed intelligence and highlights the importance of updating test norms regularly. The effect has significant implications for test interpretation and educational policy. While individual scores are still meaningful within a given time period, cross-generational comparisons require careful consideration of when tests were normed and administered.
Question 20
A counselor interprets an IQ score of 100; what does this score represent on the standard IQ scale?
- Average performance relative to the norm group, because IQ tests are typically scaled so the mean is 100 with SD 15. (correct answer)
- One standard deviation above average, because 100 is 15 points higher than the mean on the typical IQ distribution.
- High reliability, because a score of 100 indicates the test measures intelligence validly and consistently for all individuals.
- Fixed intelligence, because an IQ of 100 represents an unchangeable trait that cannot be influenced by schooling or environment.
Explanation: An IQ score of 100 represents average performance relative to the standardization sample because IQ tests are typically scaled to have a mean of 100 and standard deviation of 15. This score indicates that the individual performed at the 50th percentile - exactly average compared to others in the norm group. The score doesn't indicate high or low ability in absolute terms, but rather describes performance relative to the comparison population. Understanding this interpretation is crucial for counselors, educators, and others who use test results in decision-making. The Flynn effect shows that population averages can shift over time, making current norms important for accurate interpretation. Different theories of intelligence may suggest that a score of 100 represents average performance on the particular cognitive abilities sampled by the test, but may not reflect all aspects of human intellectual capability. This relative interpretation helps prevent both overinterpretation and underinterpretation of what IQ scores actually mean.