AutoAlign’s Sidecar Pro demonstrated superior performance in the BELLS project (Benchmarks for the Evaluation of LLM Supervision), a research initiative evaluating how effectively guardrail frameworks prevent security issues in AI applications.
Alignment Control technology
Sidecar Pro functions as a dynamic platform that iteratively improves generative AI safety, fairness, and robustness. The system uses Alignment Controls that understand user intentions expressed through natural language narratives, grammatical patterns, or Python code. It deploys custom-tuned smaller AI models working as dynamic supervision to protect against jailbreak attempts, PII leakage, stereotypical bias, and toxicity.
What was evaluated
The BELLS project assessed multiple guardrails including Lakera Guard, Langkit in both its Injection and Proactive versions, LLM Guard (Jailbreak), Nemo, and Prompt Guard. Testing employed the HF Jailbreak Prompts, Tensor Trust Data, Dan Jailbreak, Wild Jailbreak, and LMSys datasets to evaluate performance under a range of adversarial conditions.
Results
- HF Jailbreak Prompts: Sidecar Pro achieved a perfect detection rate of 1.00, matching Lakera Guard.
- Dan Jailbreak: Sidecar Pro scored 0.94, against Lakera Guard at 0.192 and Langkit at 0.67.
- Wild Jailbreak: Sidecar Pro achieved 0.78, against LLM Guard at 0.54 and Langkit at 0.24.
- LMSys: Sidecar Pro attained 0.96 in harmful content detection and 0.99 in toxic content detection.
- Dan Normal: Sidecar Pro scored 0.99 in both harmful and toxic content detection.
Across all major safety and performance categories, Sidecar Pro consistently outperformed the competing guardrails in the evaluation.



