A model’s training examples reflect choices about collection, selection, and representation. For a supervised task, labels also define the distinctions the model is being asked to learn.

Consider whether the examples cover the situations in which the system will be used. A collection can be large while leaving important cases underrepresented. Inconsistent labels can make the intended task unclear even when the inputs themselves are plentiful.

When evaluating a system, ask how the task and data relate to the real use case. The amount of data is one part of the story; its relevance, quality, and coverage are others. Those questions help connect model performance with the work it is expected to support.

Picture this situation.

Consider a collection of sample labels that covers only easy cases. An unfamiliar input may expose a boundary the examples never explained.

A second way to look.

A small trial should have a clear stopping point. Decide which uncertainty the tool can help explore and what observation would answer the next question.
A few starting points
  1. Ask which situations the examples cover.
  2. Look for a clear labeling rule.
  3. Connect the training task with the intended use.

Follow a related question

Give the note a descriptive title.

Notes you can find again

Define the rule behind each field.

A dictionary for your columns

Keep learning

Related background to continue exploring this subject.

Google: an introduction to language models NIST: AI risk management framework
Look a little closer