Skip to main content
SeekrGuard is a platform for evaluating the safety, quality, and risk of models. It gives reviewers and engineers a shared, evidence-based measure of how a model behaves on data you provide, scored with methods you define and combined into an overall risk assessment. Use SeekrGuard to determine whether a model meets your requirements for a specific use case and to compare models against one another on the same criteria.

The evaluation workflow

Most evaluations follow the sequence below, though not every step is required for every review. A quick check might end after an examination, while a formal assessment uses the full sequence.
1

Register a model

Add the models you want to evaluate from a provider. See Register and manage models.
2

Add a dataset

Upload the test inputs to evaluate against, from a file, HuggingFace, or Seekr, and map the dataset’s columns. See Add and manage datasets.
3

Choose or build evaluators

Select evaluators from the library, or create a custom LLM-as-a-Judge evaluator. See Create evaluators.
4

Run an examination

Combine the dataset, models, and evaluators into a single batch run. SeekrGuard manages the run and reports the scores. See Run examinations.
5

Review the results

Review the results across four tabs: model performance, sample responses, interactive analysis, and agentic assessment.
6

Build risk categories and a profile

Optionally, define risk categories and weight them into a risk profile to convert raw scores into an overall risk score and risk level for each model.
SeekrGuard also provides two supporting tools. Use the Chat page to interact with a model and evaluate a response inline. Use the Comparisons page to score a single response across several models at once.

Next steps

Navigate the app

The sidebar layout and the Home page.

Core concepts

The building blocks you work with in SeekrGuard.

Glossary

Definitions of the domain terms used throughout.
Last modified on August 18, 2026