跳到主要内容
知仓学习社ZHICANG

assessment-design

Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, a…

不碰外部(只输出文字)无严重或高危命中cbrock84/headcount

它会碰到什么

扫了多少2 个文本文件,8 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Assessment design

Every question measures something. The work is making sure it measures the thing you meant, because

the alternatives — reading speed, familiarity with the format, willingness to guess — are always

available and are usually easier for the student to use.

Name the inference before writing the item

An assessment is an argument: the student did this, therefore they can do that. The argument is

where assessments fail, and it fails silently.

Write down the claim first — "can decompose a two-digit number into tens and ones" — then ask what

performance would be evidence for it, and what performance would be evidence against. An item

that a student who lacks the skill can still get right is not evidence. An item that a student who

has the skill can still get wrong, for reasons unrelated to it, is worse: it produces a false

negative that gets acted on.

The most common unrelated reason is reading. Any item whose stem is harder to read than the skill is

to perform has quietly become a reading assessment.

Match the item type to the claim

  • Selected response (multiple choice, matching, true/false) is efficient and can only ever

provide evidence of recognition. A student who can recognize the correct answer cannot be assumed

to produce it.

  • Constructed response shows the path, which is what makes partial understanding visible. It

costs scoring time and needs a rubric written before the responses arrive, not after.

  • Performance tasks are the only honest evidence for anything described as applying,

investigating or designing — most NGSS performance expectations and every C3 inquiry, for

instance, cannot be assessed by selected response at all.

Mismatch is the usual failure: a standard describing explanation assessed by a multiple-choice item,

which measures whether the student can pick an explanation someone else wrote.

Distractors are the diagnostic

In a well-built multiple-choice item, each wrong answer is the result of a specific, predictable

error. Then the pattern of wrong answers says what to reteach, and the item earns its place.

Distractors that are merely wrong — a random number, an obviously absurd option — turn a four-option

item into a two-option one and tell you nothing beyond right or wrong.

The mechanical tells of a weak item, all of which students learn to exploit long before they learn

the content:

  • The longest or most qualified option is correct.
  • One option is grammatically inconsistent with the stem.
  • "All of the above" appears, and is usually correct.
  • Two options are synonyms, so neither can be right.
  • The correct answer repeats wording from the stem.

One item is not evidence

A single item carries noise — a misread word, a slip, a lucky guess — that swamps the signal for any

individual student. Inferring mastery from one response is the most common measurement error in

classroom material, and the most consequential, because it gets recorded.

Several items per claim, varied in surface form so that recognition of the format is not what is

being measured. If a claim is worth recording against a student's name, it is worth three items.

Spacing matters as much as quantity: performance on a skill immediately after it is taught measures

something closer to short-term recall than to learning. The same item a fortnight later measures

more.

Formative and summative are different products

Formative assessment exists to change what happens next, which means it must be quick, frequent,

low-stakes, and read immediately. An assessment that takes a week to score cannot be formative

whatever it is called.

Summative assessment exists to record a judgment, which means it must be defensible: enough items,

a rubric written in advance, and conditions that were the same for everyone.

Material sold as practice is usually formative in function and summative in appearance, which is

worth stating plainly to a buyer. education:learning-materials-design covers the practice

sequence this sits inside.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions

here — what each one is authoritative for, and what you may do with it. Check them before

answering on anything they cover, and cite what you used. Most are free to read and not free

to reproduce; the use note on each is binding.

Never

  • Write an item before naming the claim it is evidence for.
  • Assess an explanation standard with a recognition item.
  • Build distractors that are wrong without being diagnostic.
  • Record mastery from a single response.
  • Make the stem harder to read than the skill is to perform.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。