The evaluation workflow
Most evaluations follow the sequence below, though not every step is required for every review. A quick check might end after an examination, while a formal assessment uses the full sequence.1
Register a model
Add the models you want to evaluate from a provider. See Register and manage models.
2
Add a dataset
Upload the test inputs to evaluate against, from a file, HuggingFace, or Seekr, and map the dataset’s columns. See Add and manage datasets.
3
Choose or build evaluators
Select evaluators from the library, or create a custom LLM-as-a-Judge evaluator. See Create evaluators.
4
Run an examination
Combine the dataset, models, and evaluators into a single batch run. SeekrGuard manages the run and reports the scores. See Run examinations.
5
Review the results
Review the results across four tabs: model performance, sample responses, interactive analysis, and agentic assessment.
6
Build risk categories and a profile
Optionally, define risk categories and weight them into a risk profile to convert raw scores into an overall risk score and risk level for each model.
Next steps
Navigate the app
The sidebar layout and the Home page.
Core concepts
The building blocks you work with in SeekrGuard.
Glossary
Definitions of the domain terms used throughout.