Linguistic Distance and Translation Differential Item Functioning on Trends in International Mathematics and Science Study Mathematics Assessment Items

The 2015 Trends in International Mathematics and Science Study (TIMSS) involved 57 countries and 43 different languages to assess students’ achievement in mathematics and science. The purpose of this study is to evaluate whether items and test scores are affected as the differences between language families and cultures increase. Using differential item functioning (DIF) procedures, we compared the consistency of students’ performance across three combinations of languages and countries: (a) same language but different countries, (b) same countries but different languages, and (c) different languages and different countries. The analyses consisted of the detection of the number of DIF items for all paired comparisons within each condition, the direction of DIF, the magnitude of DIF, and the differences between test characteristic curves. As the countries were more distant with respect to cultures and language families, the presence of DIF increased. The magnitude of DIF was greatest when both language and country differed, and smallest when the languages were same, but the countries were different. Results suggest that when TIMSS results are compared across countries, the language- and country-specific differences which could reflect cultural, curriculum, or other differences should be considered.

Download Full-text

Correction of Differentially Functioning Items: Basis for Maintaining and Enhancing Test Validity and Reliability

World Journal of Educational Research ◽

10.22158/wjer.v4n1p62 ◽

2016 ◽

Vol 4 (1) ◽

pp. 62

Author(s):

Jose Q. Pedrajita

Keyword(s):

Differential Item Functioning ◽

Internal Consistency ◽

Private School ◽

Content Validity ◽

Test Scores ◽

Concurrent Validity ◽

Test Validity ◽

Internal Consistency Reliability ◽

School Type ◽

Item Functioning

This study looked into differentially functioning items in a Chemistry Achievement Test. It also examined the effect of eliminating differentially functioning items on the content and concurrent validity, and internal consistency reliability of the test. Test scores of two hundred junior high school students matched on school type were subjected to Differential Item Functioning (DIF) analysis. One hundred students came from a public school, while the other 100 were private school examinees. The descriptive-comparative research design utilizing differential item functioning analysis and validity and reliability analysis was employed. The Chi-Square, Distractor Response Analysis, Logistic Regression, and the Mantel-Haenszel Statistic were the methods used in the DIF analysis. A six-point scale ranging from inadequate to adequate was used to assess the content validity of the test. Pearson r was used in the concurrent validity analysis. The KR-20 formula was used for estimating the internal consistency reliability of the test. The findings revealed the presence of differentially functioning items between the public and private school examinees. The DIF methods differed in the number of differentially functioning items identified. However, there was a high degree of correspondence between the Logistic Regression and Mantel-Haenszel Statistic. After the elimination of the differentially functioning items, the content and the concurrent validity, and the internal consistency reliability differed per DIF method used. The content validity of the test differed ranging from slightly adequate to moderately adequate in the number of items retained. The concurrent validity of the test also differed but all were positive and indicate moderate relationship between the examinees’ test scores and their GPA in Science III. Likewise, the internal consistency reliability of the test differed. The more differentially functioning items eliminated, the lesser was the content and concurrent validity, and internal consistency reliability of the test becomes. Elimination of differentially functioning items diminishes content and concurrent validity, and internal consistency reliability, but could be use as basis in enhancing content, concurrent as well as internal consistency reliability by replacing eliminated DIF items.

Download Full-text

Determining differential item functioning and its effect on the test scores of selected pib indexes, using item response theory techniques

SA Journal of Industrial Psychology ◽

10.4102/sajip.v27i2.783 ◽

2001 ◽

Vol 27 (2) ◽

Author(s):

Pieter Schaap

Keyword(s):

Item Response Theory ◽

Differential Item Functioning ◽

Item Response ◽

Test Scores ◽

Response Theory ◽

South Africans ◽

Test Characteristics ◽

Potential Index ◽

Item Functioning

The objective of this article is to present the results of an investigation into the item and test characteristics of two tests of the Potential Index Batteries (PIB) in terms of differential item functioning (DIP) and the effect thereof on test scores of different race groups. The English Vocabulary (Index 12) and Spelling Tests (Index 22) of the PIB were analysed for white, black and coloured South Africans. Item response theory (IRT) methods were used to identify items which function differentially for white, black and coloured race groups. Opsomming Die doel van hierdie artikel is om die resultate van n ondersoek na die item- en toetseienskappe van twee PIB (Potential Index Batteries) toetse in terme van itemsydigheid en die invloed wat dit op die toetstellings van rassegroepe het, weer te gee. Die Potential Index Batteries (PIB) se Engelse Woordeskat (Index 12) en Spellingtoetse (Index 22) is ten opsigte van blanke, swart en gekleurde Suid-Afrikaners ontleed. Itemresponsteorie (IRT) is gebruik om items te identifiseer wat as sydig (DIP) vir die onderskeie rassegroepe beskou kan word.

Download Full-text

Exploring Crossing Differential Item Functioning by Gender in Mathematics Assessment

International Journal of Testing ◽

10.1080/15305058.2015.1057639 ◽

2015 ◽

Vol 15 (4) ◽

pp. 337-355 ◽

Cited By ~ 1

Author(s):

Yoke Mooi Ong ◽

Julian Williams ◽

Iasonas Lamprianou

Keyword(s):

Differential Item Functioning ◽

Mathematics Assessment ◽

Item Functioning

Download Full-text

A Comparative Study on TIMSS Mathematics Assessment of Korea, Japan, and the USA : A Differential Item Functioning Approach

Korean Comparative Education Society ◽

10.20306/kces.2017.27.5.1 ◽

2017 ◽

Vol 27 (5) ◽

pp. 1-19 ◽

Cited By ~ 1

Author(s):

Chanho Park ◽

Keyword(s):

Comparative Study ◽

Differential Item Functioning ◽

Mathematics Assessment ◽

Item Functioning ◽

The Usa

Download Full-text

Identifying Country-Specific Cultures of Physics Education: A differential item functioning approach

International Journal of Science Education ◽

10.1080/09500693.2012.684804 ◽

2012 ◽

Vol 34 (16) ◽

pp. 2483-2500 ◽

Cited By ~ 5

Author(s):

Vanes Mesic

Keyword(s):

Differential Item Functioning ◽

Physics Education ◽

Item Functioning ◽

Country Specific

Download Full-text

AN INTRODUCTION TO THE RASCH MEASUREMENT MODEL: A CASE OF MATHEMATICS EDUCATION STUDENTS COMPREHENSIVE TEST

KALAMATIKA Jurnal Pendidikan Matematika ◽

10.22236/kalamatika.vol5no1.2020pp51-60 ◽

2020 ◽

Vol 5 (1) ◽

pp. 51-60

Author(s):

Elizar Elizar ◽

Cut Khairunnisak

Keyword(s):

Mathematics Education ◽

Differential Item Functioning ◽

Rasch Analysis ◽

Item Difficulty ◽

Measurement Model ◽

Mathematical Understanding ◽

Mathematics Assessment ◽

Test Theory ◽

Item Functioning ◽

Comprehensive Test

Mathematics assessments should be designed for all students, regardless of their background or gender. Rasch analysis, developed based on Item Response Theory (IRT), is one of the primary tools to analyse the inclusiveness of mathematics assessment. However, the mathematics test development has been dominated by Classical Test Theory (CTT). This study is a preliminary study to evaluate the mathematics comprehensive test. This study aims to demonstrate the use of Rasch analysis by assessing the appropriateness of the mathematics comprehensive test to measure students' mathematical understanding. Data were collected from one cycle of mathematics comprehensive test involving 48 undergraduate students of mathematics education department. Rasch analysis was conducted using ACER Conquest 4 software to assess the item difficulty and differential item functioning (DIF). The findings show that the item related to geometry is the easiest question for students, while item concerning calculus as the hardest question. The test is viable to measure students’ mathematical understanding as it shows no evidence of Differential Item Functioning (DIF). Gender has been drawn for each of the test items. The assessment showed that the test was inclusive. More application of Rasch analysis should be conducted to create a thorough and robust mathematics assessment.

Download Full-text

Grade-Related Differential Item Functioning in General English Proficiency Test-Kids Listening

Frontiers in Psychology ◽

10.3389/fpsyg.2021.767244 ◽

2021 ◽

Vol 12 ◽

Author(s):

Linyu Liao ◽

Don Yao

Keyword(s):

Differential Item Functioning ◽

Test Scores ◽

Proficiency Test ◽

English Proficiency ◽

Language Testing ◽

Chi Square ◽

Test Fairness ◽

The Sustainable Development ◽

Item Functioning ◽

General English Proficiency Test

Differential Item Functioning (DIF) analysis is always an indispensable methodology for detecting item and test bias in the arena of language testing. This study investigated grade-related DIF in the General English Proficiency Test-Kids (GEPT-Kids) listening section. Quantitative data were test scores collected from 791 test takers (Grade 5 = 398; Grade 6 = 393) from eight Chinese-speaking cities, and qualitative data were expert judgments collected from two primary school English teachers in Guangdong province. Two R packages “difR” and “difNLR” were used to perform five types of DIF analysis (two-parameter item response theory [2PL IRT] based Lord’s chi-square and Raju’s area tests, Mantel-Haenszel [MH], logistic regression [LR], and nonlinear regression [NLR] DIF methods) on the test scores, which altogether identified 16 DIF items. ShinyItemAnalysis package was employed to draw item characteristic curves (ICCs) for the 16 items in RStudio, which presented four different types of DIF effect. Besides, two experts identified reasons or sources for the DIF effect of four items. The study, therefore, may shed some light on the sustainable development of test fairness in the field of language testing: methodologically, a mixed-methods sequential explanatory design was adopted to guide further test fairness research using flexible methods to achieve research purposes; practically, the result indicates that DIF analysis does not necessarily imply bias. Instead, it only serves as an alarm that calls test developers’ attention to further examine the appropriateness of test items.

Download Full-text

Mathematics Beliefs and Achievement of a National Sample of Native American Students: Results from the Trends in International Mathematics and Science Study (TIMSS) 2003 United States Assessment

Psychological Reports ◽

10.2466/pr0.104.2.439-446 ◽

2009 ◽

Vol 104 (2) ◽

pp. 439-446 ◽

Cited By ~ 2

Author(s):

J. Daniel House

Keyword(s):

United States ◽

Native American ◽

Test Scores ◽

National Sample ◽

Science Study ◽

Mathematics Beliefs ◽

Native American Students ◽

American Students ◽

Timss 2003 ◽

Mathematics And Science

Recent mathematics assessment findings indicate that Native American students tend to score below students of the ethnic majority. Findings suggest that students' beliefs about mathematics are significantly related to achievement outcomes. This study examined relations between self-beliefs and mathematics achievement for a national sample of 130 Grade 8 Native American students from the Trends in International Mathematics and Science Study (TIMSS) 2003 United States sample of ( M age =14.2 yr., SD = 0.5). Multiple regression indicated several significant relations of mathematics beliefs with achievement and accounted for 26.7% of the variance in test scores. Students who earned high test scores tended to hold more positive beliefs about their ability to learn mathematics quickly, while students who earned low scores expressed negative beliefs about their ability to learn new mathematics topics.

Download Full-text

Erfassung von mathematischen Kompetenzen im Vorschulalter mit MARKO-D

Diagnostica ◽

10.1026/0012-1924/a000258 ◽

2021 ◽

Vol 67 (1) ◽

pp. 13-23

Author(s):

Ariana Garrote ◽

Elisabeth Moser Opitz

Keyword(s):

Differential Item Functioning ◽

Item Functioning

Zusammenfassung. In dieser Studie wurde der Test MARKO-D (Mathematik- und Rechenkonzepte im Vorschulalter–Diagnose) mit einer Stichprobe von Kindern aus der deutschsprachigen Schweiz ( N = 555) im ersten und zweiten Kindergartenjahr erprobt und es wurde analysiert, ob sich die Altersnormen der deutschen Stichprobe auf die Schweiz übertragen lassen. Zudem wurde der Test mit einer Teilstichprobe ( n = 87) hinsichtlich Messinvarianz über die Zeit untersucht. Die Ergebnisse des eindimensionalen Rasch-Modells zeigen, dass das Instrument für die Schweiz geeignet ist. Die Testleistungen hängen jedoch vom Kindergartenbesuch ab. Für die Schweiz müssten deshalb nebst Altersnormen auch Normen pro Kindergartenhalbjahr verwendet werden. Die Analyse mittels Differential Item Functioning ergab, dass 17 von 55 Items von großer Messvarianz über die Zeit betroffen sind. Um das Instrument für Längsschnittuntersuchungen einsetzen zu können, müsste es weiterentwickelt werden.

Download Full-text

Differential Item Functioning in Brief Instruments of Disordered Eating

European Journal of Psychological Assessment ◽

10.1027/1015-5759/a000472 ◽

2019 ◽

Vol 35 (6) ◽

pp. 823-833 ◽

Cited By ~ 4

Author(s):

Desiree Thielemann ◽

Felicitas Richter ◽

Bernd Strauss ◽

Elmar Braehler ◽

Uwe Altmann ◽

...

Keyword(s):

Differential Item Functioning ◽

Disordered Eating ◽

Structural Equation ◽

Young Female ◽

Eating Attitudes ◽

Equation Model ◽

German Population ◽

Test Fairness ◽

Item Functioning ◽

Multiple Indicator

Abstract. Most instruments for the assessment of disordered eating were developed and validated in young female samples. However, they are often used in heterogeneous general population samples. Therefore, brief instruments of disordered eating should assess the severity of disordered eating equally well between individuals with different gender, age, body mass index (BMI), and socioeconomic status (SES). Differential item functioning (DIF) of two brief instruments of disordered eating (SCOFF, Eating Attitudes Test [EAT-8]) was modeled in a representative sample of the German population ( N = 2,527) using a multigroup item response theory (IRT) and a multiple-indicator multiple-cause (MIMIC) structural equation model (SEM) approach. No DIF by age was found in both questionnaires. Three items of the EAT-8 showed DIF across gender, indicating that females are more likely to agree than males, given the same severity of disordered eating. One item of the EAT-8 revealed slight DIF by BMI. DIF with respect to the SCOFF seemed to be negligible. Both questionnaires are equally fair across people with different age and SES. The DIF by gender that we found with respect to the EAT-8 as screening instrument may be also reflected in the use of different cutoff values for men and women. In general, both brief instruments assessing disordered eating revealed their strengths and limitations concerning test fairness for different groups.

Download Full-text