Article

Sidecar Pro outperforms other AI security guardrails in BELLS research

An independent benchmark of guardrail frameworks put Sidecar Pro ahead of every other system tested, across five adversarial datasets.

· Rahm Hafiz, CTO and co-founder, AutoAlign

Stylised bar chart with one lime bar rising clearly above the rest, crowned by a shield

AutoAlign’s Sidecar Pro demonstrated superior performance in the BELLS project (Benchmarks for the Evaluation of LLM Supervision), a research initiative evaluating how effectively guardrail frameworks prevent security issues in AI applications.

Alignment Control technology

Sidecar Pro functions as a dynamic platform that iteratively improves generative AI safety, fairness, and robustness. The system uses Alignment Controls that understand user intentions expressed through natural language narratives, grammatical patterns, or Python code. It deploys custom-tuned smaller AI models working as dynamic supervision to protect against jailbreak attempts, PII leakage, stereotypical bias, and toxicity.

What was evaluated

The BELLS project assessed multiple guardrails including Lakera Guard, Langkit in both its Injection and Proactive versions, LLM Guard (Jailbreak), Nemo, and Prompt Guard. Testing employed the HF Jailbreak Prompts, Tensor Trust Data, Dan Jailbreak, Wild Jailbreak, and LMSys datasets to evaluate performance under a range of adversarial conditions.

Results

  • HF Jailbreak Prompts: Sidecar Pro achieved a perfect detection rate of 1.00, matching Lakera Guard.
  • Dan Jailbreak: Sidecar Pro scored 0.94, against Lakera Guard at 0.192 and Langkit at 0.67.
  • Wild Jailbreak: Sidecar Pro achieved 0.78, against LLM Guard at 0.54 and Langkit at 0.24.
  • LMSys: Sidecar Pro attained 0.96 in harmful content detection and 0.99 in toxic content detection.
  • Dan Normal: Sidecar Pro scored 0.99 in both harmful and toxic content detection.

Across all major safety and performance categories, Sidecar Pro consistently outperformed the competing guardrails in the evaluation.

Get started

Ready to talk?

If operational intelligence matters to your production line, your field service operation, or your programme, we want to talk.