Skip to main content
SeekrGuard measures how a model behaves on your data, using criteria you define. A model, a dataset, and an evaluator combine into an examination, and its scores roll up through risk categories into a profile that rates the model’s risk. For the precise definition of any term, see the glossary.

Provide inputs

An evaluation draws on three inputs: These inputs depend on one another. Each evaluator declares the dataset fields it needs, such as a user question, a context passage, or an expected answer, so it can run only on a dataset that supplies those fields and against outputs a model has produced.

Run evaluations

You put the inputs to work in two ways:
  • Examination – a batch run that scores one or more models across all or part of a dataset and reports aggregated results. See Run examinations.
  • Comparison – a focused check that scores a single response with several evaluators and judge models at once. See Run comparisons.
  • Chat – talk to a model directly and score any response on the spot.
Use an examination for systematic vetting, a comparison for a quick side-by-side look, and chat for probing a model as you go. Only examination results are stored for risk scoring. Comparison and chat scores are read where they appear and go no further.

Assess risks

Individual evaluator scores become useful once they combine into a single, weighted measure of risk: For exactly how scores are normalized, weighted, and combined, see how a risk score is calculated.
Last modified on August 18, 2026