Showing posts with label IRT. Show all posts
Showing posts with label IRT. Show all posts

Monday, July 1, 2019

Gathering Rich Data to Support Effective Teaching and Learning

Information about a student’s capabilities gathered through the assessment and instructional process creates a portrait of the student as a learner. This type of feedback for both teacher and student will set the stage for establishing goals for improving student learning. Instruction can be strategically planned and lessons can be identified to target students at a variety of levels. 

Understanding what students in a classroom already know is the first step. The Galileo Benchmarks series includes three assessments that can be administered at the beginning, middle, and end of year. The first assessment administered will provide a placement for students. These assessments provide an IRT scale score as well as information about student mastery of individual domains. The two additional benchmark tests are given in intervals throughout the year tracking progress. From these benchmarks, data points identify which students are at-risk and which students are meeting benchmark goals. This is illustrated through changes in the student performance levels and risk levels.

Galileo Risk Level Summary Report
Risk Levels used to inform district decision-making 


This rich data is available within the Galileo comprehensive reporting suite. The Dashboards reports support the continuous gathering of data to support effective teaching and learning, resulting in high student achievement that is based on today’s standards for college and career readiness. Educators can use Galileo to track student growth and achievement, which includes evaluating student progress toward standards mastery as well as college and career readiness as measured by statewide assessments. 

The reports also provide detailed information about student strengths and weaknesses including their level of mastery of individual state standards. This powerful portrait of student learning is made possible by transforming high-powered data analysis procedures such as IRT, test scaling procedures, and inferential statistics into practical, interactive reports that educators can use to facilitate quality instruction, enrichment and re-teaching activities. The reports present multiple measures of student performance, including raw scores (number/percent correct), Developmental Level (IRT scale) score ability estimates, standards mastery classifications, growth, and risk information.  

Learn more about the real-time reporting functions in the Galileo comprehensive reporting suite. Check out the website or contact us for a personalize demonstration. 

Monday, November 5, 2018

Experience the Ways the Galileo Comprehensive Assessment System Promotes Student Learning—Complimentary Trial Available Now

Contact us to start today.
Assessment is a critical part of effective instruction. Research has shown that educators who implement reliable and valid assessments to guide instructional decisions will promote student learning. That’s why ATI’s Galileo K-12 comprehensive assessment system is designed to help inform educational decisions promoting student learning in K-12 math, ELA, science, and other subjects. 

See it in action and contact us for a Free Trial! 

What makes the comprehensive assessment system exceptional?
Let’s begin with its broad array of standards-based, curriculum and pacing guide aligned customized assessments built to client specifications every year. ATI’s in-house Educational Management Services and Research teams, based on your assessment requests, do the assessment construction work for you. The resulting tests will not have accidental duplications across the academic year, will be free of item enemies, will address the Depth of Knowledge levels you specify, and will contain, if you so desire, technology enhanced items.

An additional option provided through the system is a wide variety of high-quality, pre-built assessments. These include benchmark, formative, pretest and posttest, placement, summative, computer adaptive, end-of-course, college prep assessments, plus more.  

ATI’s Secure and Community Item Banks offer 19 item types including constructed response and technology-enhanced items. Each of these item types is designed to engage students in complex thinking to help ensure college and career readiness. Educators can also access more assessment and item choices through the integrated Inspect® and Certica item banks, including Certica’s Navigate Item Bank™.

ATI understands that educators also need to be able to create their own items and assessments quickly and easily. To facilitate this, our intuitive and powerful Test Builder and Item Builder interfaces make creating your own assessments and items a snap. 

Taking a look under the hood—Item Response Theory 
Options are great, but to understand what really makes the Galileo comprehensive assessment system so special, we need to take a look under the hood. Powered by ATI research, Galileo takes advantage of Item Response Theory (IRT), and other state-of-the-art statistical analyses to go beyond systems that only report raw scores and percent correct. ATI uses IRT to measure the Developmental Level (DL) of each student. The DL score indicates what students have already mastered and what they are ready to learn next.

Galileo assessments place DL scores and measures of test difficulty on a common scale ensuring that DL score changes reflect growth, not variations in test difficulty. DL scores can then be used to effectively forecast student performance on statewide assessments. Galileo’s innovative Dashboards graphically illustrate this critical information to help educators apply their efforts where needed most—making instruction even more effective.

Try the comprehensive assessment system through a complimentary trial 
Contact us to start your complimentary trial of the Galileo comprehensive assessment system. With high-powered data, easy to use tools, and robust item banks, the Galileo comprehensive assessment system is designed to promote success for all students. This is the assessment system you've been searching for. Own your assessment-process and let research work for you, with the Galileo K-12 comprehensive assessment system.   
 
To learn more about other Galileo services and features, visit us at




Monday, October 8, 2018

Educators Benefit from ATI's Research and New IRT Item Parameters

Did you know that each year ATI refreshes and maintains stable estimated parameters for items in the ATI Secure and Community Item Banks based on tens of thousands of students? ATI uses item parameters to help build reliable, valid district/charter-wide assessments, and to evaluate item performance to guide item bank revisions.




Test Builder powerful tools support the creation of reliable and valid assessments.

You might ask why yearly refreshment matters. It matters because the quality of item parameter estimates has a direct effect on the quality of the test therefore on the quality of the data from the tests. ATI uses item parameters to build reliable, valid district/charter-wide assessments and to evaluate item performance to guide item bank revisions. Quality of item parameter estimates also impact the accuracy of student ability estimates. Accurate student ability estimates provide the ability to compare performance on a variety of assessments covering the same or different content areas and to measure academic progress across time. 

Galileo puts the power of Item Response Theory (IRT) in the hands of educators. Educators have always been able to view item parameters for district/charter-wide assessments in the Test Review interface and in the Item Parameters Report

Now, educators can also view item parameters along with other metadata in Test Builder and in drill downs from many Galileo reports such as Test Monitoring, Intervention Alert, Item Analysis, and Detailed Benchmark Performance Levels.

Check out these new features to take advantage of valuable IRT parameter information as you are creating tests and evaluating student performance to guide instruction! Learn more by accessing a short video on “What Role Does IRT Play in Assessments?” Contact a Field Services Coordinator to learn more.

Other resources:
Make a Measurable Difference with ATI and IRT Research
What Makes the ATI Item Banks the Preferred Choice by So Many Educators?
Experience the Ways the Galileo Comprehensive Assessment System Promotes Student Learning — Complimentary Trial Available

Monday, October 30, 2017

Comparing Results from Benchmarks When the Questions and Standards Are Different

The administration of benchmark assessments is a common practice with which most educators are familiar. Benchmarks assessments, sometimes referred to as interim assessments, are intended to measure a student’s mastery of skills in a content area against grade-level standards and learning goals. But what happens when the assessments vary? What do you do when the questions and standards are not the same across benchmark assessments? How do you evaluate student progress when the benchmark assessments are not identical to one another?

The team of researchers at ATI make it possible to answer these questions by implementing state-of-the-art measurement techniques. The process begins by using Item Response Theory (IRT) scaling techniques to place student scores on a series of Galileo assessments on a common scale. The relationship between student ability and item difficulty plays a particularly important role in the IRT scaling process. In IRT, student ability estimates and item difficulty estimates inform one another within the context of the same mathematical model. Student ability is estimated in light of the relative difficulty of the items on the test, and the difficulty of the items is estimated in light of the ability level of the students who responded to them. For example, in a common IRT model, a student of average ability will have a fifty-fifty chance of responding correctly to an item of average difficulty. A student who is one standard deviation above the mean ability level will have a fifty-fifty chance of responding correctly to an item that is one standard deviation above the mean in terms of difficulty. Likewise, a student of below average ability will have a fifty-fifty chance of responding correctly to a corresponding item that is below average in difficulty. 

The fact that ability and difficulty are measured on the same scale makes it possible to adjust the student’s scale score, which is an estimate of ability, based on the difficulty of the items included in the assessment. This adjustment is a key factor in the scaling process making it possible to compare scores from different tests. When scores such as percent correct are used, such adjustment is not possible, and scores from different tests cannot be compared. For example, if a student received a score of 70 percent correct on one test and 90 percent correct on a second test, the difference could have occurred because the second test was easier than the first, or because of an increase in student performance, or both.  

In this regard, a powerful report to check out in the Galileo Help files is the Item Parameters Report. The report specifically provides information about item difficulty and other item parameters for each item on a benchmark test. Other helpful reports to check out are the Aggregate Multi-Test and Student Growth and Achievement reports, which present student IRT scale scores, which are called Developmental Level, or DL scores in Galileo, on a series of tests so that student progress may be monitored.

Watch the brief video on Psychometrics 

To learn more about ATI’s research initiatives, visit our website. For a one-on-one demonstration of the reporting features mentioned in this blog, request a personal demo with one of our knowledgeable field services coordinators. 

Other topics of interest:
How to use item parameters to make decisions during test review
How does ATI calculate my district’s psychometrics benchmark test data?

Monday, February 22, 2016

New and Upcoming Enhancements to Galileo

New – Depth of Knowledge Level Filters in Test Builder and Displays in Reports  
ATI has made it easier for districts and charters to develop assessments addressing appropriate levels of cognitive complexity. Galileo Test Builder has been enhanced to support filtering item searches by Depth of Knowledge (DOK) level while developing assessments. Additionally, when items are administered, DOK information may be easily viewed in several reports (i.e. Item Analysis, Intervention Alert, and Test Monitoring Reports). Soon, DOK levels will also be viewable in the Test Review interface for users with appropriate permissions.  


Filter item searches by one, multiple, or all DOK levels. Click to view larger. 
Technology enhanced DOK level 3 item displayed in the drilldown item view of the Intervention Alert Report.  Click to view larger.
Learn more about searching for items in Test Builder by viewing the Quick Reference Guide.

Coming Soon – Real-Time Developmental Level Scores and Statewide Test Forecasting
Galileo will begin automatically generating Developmental Level (DL) scores, performance levels, and risk levels for Arizona students as soon as they have completed district/charter-wide assessments. The DL score is an Item Response Theory (IRT) based measure of student ability. The DL score takes into account the difficulty of test items and can be placed on a common scale across tests to evaluate growth. The DL score is also used to classify students into Risk Levels that forecast likely student performance on the statewide test at the end of the year. Immediate access to DL scores and Risk Levels will provide districts/charters with real-time actionable information to guide next instructional steps. 

For more information on these or other features of Galileo, contact our friendly and knowledgeable Field Services Coordinators

Tuesday, July 21, 2015

No Crystal Ball Needed: New Resource Tools Help Early Childhood Providers Evaluate Children’s Learning

Galileo Pre-K Online is a reliable and valid assessment tool that can be used to define, assess, and track learning. Galileo applies procedures based in Item Response Theory (IRT) to information gained through observational assessment to estimate a measure of child learning for each age range scale. This estimate of child learning is presented as the Developmental Level (DL) score.

ATI conducts research and analysis each year creating resources to help evaluate children’s learning. Two DL score analysis research briefs are now available:

Predicted Developmental Level Scores for Children Birth through 5 Years Based on 2014-15 Assessment Data
The predicted DL score helps you evaluate your children’s learning at various ages relative to the predicted DL score for children that age. Having this information early in the program year means you can identify children whose DL score is higher or lower than predicted and provide them enrichment and or intervention activities to further their development.


Estimated Growth for Children Birth through 5 Years Based on 2014-15 Assessment Data
In Galileo, growth is defined as a child’s change in his or her DL score. To help determine the estimated growth of children over various time periods, ATI looked at the relationship between a child’s DL score and time (in days). From this analysis, ATI was able to determine the estimated daily growth rate.


Read the full research briefs:
• Predicted Developmental Level Scores for Children Birth through 5 Years
• Estimated Growth for Children Birth through 5 Years


Monday, April 20, 2015

How to Use Item Parameters to Make Decisions During Test Review

Assessment Technology’s use of Item Response Theory (IRT) provides clients with rich information about items. This information can be available for items written by district professionals as well as those written by ATI item writers.

Explanations of the provided data points is provided below.

Parameter Definitions
Understanding IRT Item Parameters
The best way to understand what item parameters refer to is to look at an Item Characteristic Curve. On an item characteristic curve, which presents the data for one, specific item, student ability (based on their performance on the assessment as a whole) is plotted on the horizontal axis, with a mean of 0 and a standard deviation of 1. The probability of answering the item correctly is plotted on the vertical axis. Typically the probability of answering correctly is relatively low for low ability students and relatively high for students of higher ability.


Difficulty Parameter
The first example (Item 4) is an example of a great item as far as the parameters go. The b-value (difficulty) for that item was 0.689, which is a bit on the difficult side, but not too bad. The important point here is that b-values (item difficulty) are on the same scale as student ability. So what this example is telling us is that students at or above 0.689 standard deviations above the mean are likely to get the answer correct. Students below that point on the ability scale are more likely to answer incorrectly. The b-parameter is also known as the location parameter, because it locates the point on the ability scale where students start demonstrating mastery of the concept.

 
Figure 1
Item 4 with a difficult b-value.


Tip: This parameter should have a wide range, generally between -3 and +3, across a test.

Discrimination Parameter
The a-value (discrimination) refers to how well the item discriminates between different ability levels. It’s how steep the rise is in the curve that shows the probability of answering correctly. Ideally, there is a nice, steep rise in the probability of answering correctly like the one for test Item 4. That indicates that there is a dramatic change in how likely it is that a student has mastered the concept that’s pin-pointed within a very narrow range of the ability scale. You can be pretty confident that students above 0.689 standard deviations above the mean “get it” and that students below that point generally don’t. The discrimination parameter for test item 4 is 1.459.

The next example, Item 5, shows an item that doesn’t discriminate quite as well as Item 4. The a-value on that one is 0.53. It’s also a pretty easy item, with a b-value of -1.07. So, on this one, most students are likely to get it correct, unless they’re more than one standard deviation below the mean of the ability scale.
 
Figure 2
Item 5 with a lower b-value.


It’s not necessarily the case that an easy item automatically has poor discrimination. The final example, Item 8, is an easy item that discriminates very well. Although students at most ability levels are likely to select the correct answer, there is still a dramatic increase in likelihood of answering correctly within a relatively narrow range of the ability scale.
 

Figure 3
Item 8, is an easy item that discriminates well.


Tip:  This parameter should be near 1 or above.

Guessing Parameter
The guessing parameter (the c-value) is the probability of getting the item correct by just guessing. It defines the lower limit of the item characteristic curve. For a multiple choice item with four answer choices, the guessing parameter should be around 0.25 or, preferably, a bit lower.


Tip:  This parameter should be .25 or below for a four alternative multiple-choice test item.

As the Director of Educational Management Services, I have spent the past nine years working with teachers and reviewers in creating assessments using this information provided by IRT. I have found that this information although important is not the only consideration I use when I complete a test review. I find that the best use of item parameters is to be informed about what the data means, but at the same time, use knowledge of a district’s students, teachers, and curriculum to pick the best items to suit a specific population’s needs.

How do I use item parameters to inform my decision in test review? I believe that the best use of this data is to inform opinions and choices of a test reviewer. Let me give you an example.

I receive a comment from an initial review that says the item is too difficult for students. I check the b-value provided on the Test Review page. The item’s difficulty is a 3.00 or above, this shows that the reviewer’s intuition about the item is correct. I look to see what the b-values of the other items are on this assessment. If I find that there are alternative items with high b-values already on this test, I may decide to replace the item and place an easier item on the test.

On the other hand, if I check the b-value and find that the item has a b-value closer to 1.00, I know that other students have handled this item without a lot of difficulty. In this case, I may decide to leave the item on the assessment and see how students do on this item.

In other words, I believe that reviewer opinions and data should be used equally to inform decisions during test review.

Karyn White                                                        
Director of Educational Management Services
   

Monday, February 16, 2015

Extensive Statistical Analyses Sets Galileo K-12 Online Apart

ATI research uses advanced statistical procedures to address assessment goals associated with standards-based education. One research focus includes statistical analyses of items and assessments. Item Response Theory (IRT) a statistical analysis procedure, makes it possible to place scores from different assessments on a common scale. If IRT were not used, a difference in scores between two tests could be due to differences in test difficulty or changes in student achievement, or both. IRT considers both item difficulty and achievement level making it possible to produce an ability score and to measure academic progress. This cannot be accomplished by assessments that provide only raw data (percent/number correct).

ATI routinely provides psychometric analyses of ATI and district developed assessments that have been administered to a group of students large enough to support such analyses. Item analysis uses IRT techniques to determine item parameter estimates and assessment reliability. Access to a variety of statistical information is available through the Galileo K-12 Online Item Analysis and Item Parameter reports.

ATI also utilizes categorical data analysis procedures to determine the risk of not meeting standards. Categorical data analysis techniques informed development of the Galileo Risk Assessment Report. The report classifies students with respect to their risk of failing to show mastery on the statewide test based on their performance on one or more district -wide assessments. Structural equation modeling  is also used to identify variables that impact learning. Investigations using structural equation modeling have revealed the effects of student learning at different points throughout the year on performance on the statewide test. Finally, value-added modeling is used to identify factors affecting learning.

For more information on how ATI can support your district or charter with a comprehensive instructional improvement and instructional effectiveness system designed to promote student learning, contact us today.

Thursday, January 7, 2010

So what are item parameters, anyway?

Item Response Theory (IRT) item parameter estimates play an important role in many of the reports generated in Galileo K-12 Online. In addition to describing item characteristics in the Item Parameter report, they help to determine student Development Level (DL) scores and to identify the standards that are recommended for re-teaching efforts in the drill-downs from the Risk Assessment report. But what are they?

The best way to understand what item parameters refer to is to look at an Item Characteristic Curve. On an item characteristic curve, which presents the data for one, specific item, student ability (based on their performance on the assessment as a whole) is plotted on the horizontal axis, with a mean of 0 and a standard deviation of 1. The probability of answering the item correctly is plotted on the vertical axis. Typically the probability of answering correctly is relatively low for low ability students and relatively high for students of higher ability.

The first example (Item #4) is an example of a great item as far as the parameters go. The b-value (difficulty) for that item was 0.689, which is a bit on the difficult side, but not too bad. The important point here is that b-values (item difficulty) are on the same scale as student ability. So what this example is telling us is that students at or above 0.689 standard deviations above the mean are likely to get the answer correct. Students below that point on the ability scale are more likely to answer incorrectly. The b-parameter is also known as the location parameter, because it locates the point on the ability scale where students start demonstrating mastery of the concept.



The a-value (discrimination) refers to how well the item discriminates between different ability levels. It’s how steep the rise is in the curve that shows the probability of answering correctly. Ideally, there is a nice, steep rise in the probability of answering correctly like the one for question 4. That indicates that there is a dramatic change in how likely it is that a student has mastered the concept that’s pin-pointed within a very narrow range of the ability scale. You can be pretty confident that students above 0.689 standard deviations above the mean “get it” and that students below that point generally don’t. The discrimination parameter for question 4 is 1.459.

The next example, Item 5, shows an item that doesn’t discriminate quite as well as Item 4. The a-value on that one is 0.53. It’s also a pretty easy item, with a b-value of -1.07. So, on this one, most students are likely to get it correct, unless they’re more than one standard deviation below the mean of the ability scale.


It’s not necessarily the case that an easy item automatically has poor discrimination. The final example, Item 8, is an easy item that discriminates very well. Although students at most ability levels are likely to select the correct answer, there is still a dramatic increase in likelihood of answering correctly within a relatively narrow range of the ability scale.


The guessing parameter (the c-value) is the probability of getting the item correct by just guessing. It defines the lower limit of the item characteristic curve. For a multiple choice item with four answer choices, the guessing parameter should be around 0.25 or, preferably, a bit lower.