Community
Who We Are
cotrust.ai is a global community of researchers, engineers, product leaders, linguists, and domain experts, part of a 20,000+ member ecosystem spanning AI communities worldwide. We evaluate the latest AI models and agents for quality, safety, and instruction adherence, bringing together human experts and an open community with diverse backgrounds, skills, and perspectives to probe these systems and understand their capabilities and limitations across multimodal reasoning, tool use, and real-world tasks.
What We Do
We build the artifacts that make AI evaluation concrete: rubrics, curated datasets, open specifications, and shared frameworks, all grounded in hands-on experience with production systems. Our work spans how models and agents behave, how humans stay meaningfully in the loop, and how failures get caught.
We work in the open. Our datasets, tools, and methodologies are built to be reused, extended, and improved by the broader community.
Why It's Collective
It takes diverse expertise, adversarial thinking, and evaluation data from people who actually work with these systems to understand what they can and can't do. That's the work we do together.
Open Evals
Collect structured human judgments in the Model Playground and share evaluation datasets with the community.
While automated metrics are valuable, Human Evaluation is a crucial complementary workflow. Our Model Playground allows you to collect structured human judgments on model outputs, helping to improve model performance and contribute to open research.
1. Choose Your Evaluation Mode
From the Playground, click on Evaluate and select either Direct or Compare mode:
- Direct Mode: Rate a single model's response using thumbs-up/down and customizable label pairs (e.g., quality, safety, or instruction following). You can also apply system instructions and add RAG context here.
- Compare Mode: View two models side-by-side. You can optionally apply different system instructions to each model to evaluate how they handle the same prompt.
2. Prompt and Evaluate
Submit your prompts just as you normally would.
- In Direct mode, rate each response before moving on to the next turn.
- In Compare mode, both models stream their answers simultaneously. Once finished, choose your preference (A, B, Both Good, or Both Bad).
3. Record Your Preferences
Once finished, you can create and upload your evaluation as a structured dataset that includes prompt-response pairs, evaluation ratings and labels and metadata. Datasets marked as public will appear in the Browse community catalog, while private datasets remain exclusively on your account.
4. Manage Privacy & Open Research
When starting an evaluation session, you can choose whether to contribute your data to open research. This aligns with transparent benchmarking: human choices provide the critical signal needed to refine AI models. You can make this decision per session or set a default preference under Settings → Profile → Open Research.
Logos
Loading…
Polyglot
ContributeLoading…
Loading…
