← Open Source
HowieHwong

TrustLLM

[ICML 2024] TrustLLM: LLM trustworthiness evaluation across truthfulness, safety, fairness, robustness, privacy and ethics. Python/CLI toolkit for local models, compatible APIs and AI agent orchestration.

AI EngineeringTest & guardPython
Open on GitHub
Momentum
+0stars in 24 hours0.0%
633
Stars
68
Forks
+0
This week
13
Contributors
Created 2023-12-23 · Updated 2026-10-04 · #13699 today
Top developers
README

TrustLLM — Trustworthiness in Large Language Models

ICML 2024   ·   TRUSTWORTHINESS IN LARGE LANGUAGE MODELS

Paper   /   Documentation   /   Dataset   /   Leaderboard

English   /   简体中文   /   繁體中文   /   日本語   /   한국어   /   Español   /   Français

TrustLLM is an open research toolkit for evaluating the trustworthiness of large language models across six dimensions. Run the ICML 2024 benchmark with local weights or a model API, and keep the data, settings, and results together.

Automating evaluations with an AI agent? See the agent integration guide for CLI/Python orchestration and JSON results, or start from the machine-readable documentation index.

Start here

1 — Install. The base package supports API generation and data downloads; no GPU libraries required.

python -m pip install "trustllm @ git+https://github.com/HowieHwong/TrustLLM.git@main#subdirectory=trustllm_pkg"

These commands install the 0.4 source version directly from GitHub. Git is required. For reproducible runs, replace main with a commit SHA.

2 — Get the benchmark.

python -m trustllm download --output data
python -m trustllm tasks

3 — Try your API model. Set the credentials and endpoint for your service:

export OPENAI_API_KEY="your-api-key"
export OPENAI_BASE_URL="https://your-provider.example/v1"

python -m trustllm generate \
  --backend api --model your-model-id \
  --task safety --data data/dataset \
  --limit 5 --concurrency 4 --output runs/api-safety

Use your provider's actual API root in place of the example URL. For a local OpenAI-compatible server, use http://localhost:8000/v1 and its served model ID; an API key is optional if the server does not require one. API mode uses text-only Chat Completions.

Open runs/api-safety/report.html to inspect completion counts. --limit 5 selects five samples per dataset file for a smoke test; remove it and choose a new output directory for a full run. This report tracks generation, not benchmark scores.

Local weights, same workflow

python -m pip install "trustllm[local] @ git+https://github.com/HowieHwong/TrustLLM.git@main#subdirectory=trustllm_pkg"

python -m trustllm generate \
  --backend local --model Qwen/Qwen2.5-0.5B-Instruct \
  --task safety --data data/dataset \
  --device auto --limit 5 --output runs/local-safety

Hugging Face weights download on first use. Replace the model ID with /path/to/checkpoint to use existing weights. --device cpu, cuda:0, and mps select a device; auto uses Accelerate placement. Local loading supports Transformers causal language models and uses the tokenizer's chat template when available. Model access, hardware capacity, and architecture compatibility still apply.

Prefer Python?

from trustllm import download_dataset, generate

download_dataset("data")

run = generate(
    model="your-model-id",
    backend="api",                    # Switch to "local" for HF weights.
    task="safety",
    data_path="data/dataset",
    output_dir="runs/python-safety",
    limit=5,
)
print(run["status"], run["successful"], run["total"])

API settings come from OPENAI_API_KEY and OPENAI_BASE_URL, or explicit api_key and base_url Python arguments. See usage and configuration for JSON configs, retries, model revisions, token settings, and resuming a run.

From responses to scores

Install the scoring dependencies, then evaluate a complete directory of generated responses:

python -m pip install "trustllm[eval] @ git+https://github.com/HowieHwong/TrustLLM.git@main#subdirectory=trustllm_pkg"
python -m trustllm evaluate --task safety --data runs/api-safety-full

Replace the path with your full-run output directory. Scores are written to scores.json and scores.html; incomplete responses are rejected. The six scoring pipelines retain the original benchmark methods. Depending on the task, scoring uses rules, a downloaded classifier, embeddings, or an API judge. Set OPENAI_JUDGE_MODEL for a judge available to your account. Judge calls can incur charges. Read the scoring guide before a full run.

Dimension Evaluation focus
Truthfulness Misinformation · hallucination · sycophancy
Safety Jailbreaks · misuse · exaggerated safety
Fairness Stereotypes · preferences · disparagement
Robustness Adversarial perturbations · out-of-domain inputs
Privacy Privacy awareness · information leakage
Ethics Moral judgments · moral choices

Dataset and metric reference →

What changed in 0.4?

  • Arbitrary model IDs and a shared runner for API/local generation; no legacy model whitelist.
  • Lightweight API install; optional local, eval, and archived legacy dependencies.
  • download, tasks, generate, and evaluate commands, also through python -m trustllm.
  • Bounded API retries, per-sample checkpoints, explicit failures, and guarded --resume.
  • JSON response files remain compatible with the existing res-based evaluators.
  • Saved dataset hashes, settings, dependency versions, usage where returned, and HTML generation reports.

The original generation engine is archived under trustllm.generation.legacy. Generation formatting and model integrations have changed: new runs are not automatically equivalent to the original paper's settings. See migration notes.

Research & development

CI checks · Contributing · Design references · Changelog · Issues

The language links above translate this README. Detailed guides are currently in English; translations do not change benchmark prompts, datasets, or scoring methods.

If TrustLLM supports your research, please cite the ICML 2024 paper. Full BibTeX: CITATION.bib. Code: MIT. Dataset terms remain with the original sources.