# LangSmith Evaluations: LLM & AI Agent Evaluation Platform

## Continuously improve agent quality
Run evals before and after shipping, gather expert feedback on agent performance, and iterate on prompts with your team.

## Helping top teams ship great agents

### Three ways to build, from full harness to full control
Each package solves a different problem. Choose where you want to start, then compose the rest of the stack around it.

### Evaluate your agent’s performance
Run evaluations on curated datasets during development to compare agent versions, benchmark performance, and catch regressions before users do. Monitor performance in production with online evals that score user interactions with your agent in real-time to detect issues and measure quality.
- _Calibrate llm-as-judge evals with human feedback_
- _Conversation evals_
- _Multi-modal evals_

[Learn about eval techniques](https://docs.langchain.com/langsmith/evaluation-concepts)

### Gather expert feedback
Equip subject-matter experts to assess response quality or review specific attributes of your agent. Automatically assign runs for review, and annotate any part of your agent workflow to capture precise feedback.
- _Embedded renderings of UI in review flow_
- _Auto send interesting traces for human review_
- _Shared scoring criteria to standardize reviewer feedback_

[Streamline human feedback](https://docs.langchain.com/langsmith/set-up-feedback-criteria)

### Iterate & collaborate on prompts
Experiment with prompts in the Playground, and compare outputs across different prompt versions or model providers. Use our AI LangSmith chat to auto improve prompts. Scale spot checking by running an evaluation of the prompt on a larger dataset, all from the UI.

[Create and test a prompt](https://docs.langchain.com/langsmith/prompt-engineering-quickstart)

### Resources for LangSmith Evaluation

- [Evals guidebook](/content/testing-guide-ebook/index.html)
- [Evals concepts](https://docs.langchain.com/langsmith/evaluation-concepts)
- [Foundational course](https://academy.langchain.com/courses/intro-to-langsmith)

### FAQs for LangSmith Evaluation

**What kind of evaluators does LangSmith support?**
LangSmith's evaluation framework supports multiple evaluator types: human evaluation through annotation queues, heuristic checks (like validating outputs or checking if code compiles), LLM-as-judge evaluators that score against criteria you define, and pairwise comparisons. You can also write custom evaluators in Python or TypeScript with any business logic you need.

**How does human feedback and annotation work?**
LangSmith makes it easy for AI teams to collect expert feedback through annotation queues. Flag runs for review, assign them to subject-matter experts, and use that feedback to calibrate automated evaluation, improve prompts, or augment datasets with high-quality test cases.

**How reliable is LLM-as-judge, and how do I audit it?**
LLM-as-judge evaluators don't always get it right. LangSmith lets you route samples to human reviewers who flag disagreements, helping you identify failure modes and edge cases.

**What's the difference between offline and online evaluation?**
Offline evaluation runs against curated datasets during development to catch regressions before deployment. They act as unit tests for your LLM application.

**Can I use LangSmith Evaluation without LangSmith Observability?**
Yes. You can use LangSmith Evaluation with or without Observability. For all plan types, you'll get access to both and only pay for what you use.

**How does LangSmith evaluate AI agents and multi-turn workflows?**
Agent evaluation in LangSmith captures the full trajectory of steps, tool calls, and reasoning your agent took.

**Can I run evaluations in my CI/CD pipeline?**
Yes. LangSmith integrates with pytest, Vitest, and GitHub workflows so you can run evals on every PR or nightly build.

**How do I benchmark across prompts, models, or agent versions?**
LangSmith's comparison view dashboards show results side-by-side across experiments.

**Do I have to use LangChain or LangGraph to use LangSmith?**
No. LangSmith is framework-agnostic. Evaluate AI applications built with LangGraph, custom Python, or any other framework.

**How do I get started if I don't have a labeled dataset?**
Start by capturing production traces with LangSmith, then sample interesting or problematic runs into a dataset.

**Will LangSmith add latency to my application?**
No. The LangSmith SDK uses an async callback handler that sends traces to a distributed collector.

**Can I self-host LangSmith? Where is my data stored?**
LangSmith instances hosted at [smith.langchain.com](https://smith.langchain.com/) stores data in GCP us-central-1 or europe-west4.

**Will you train on the data that I send LangSmith?**
We will not train on your data, and you own all rights to your data.

**How much does LangSmith cost?**
LangSmith has a free tier for development and small-scale production. Paid plans scale with trace volume.

### Ready to build better agents through continuous evaluation?
[Start building](https://smith.langchain.com/) [Get a demo](/content/contact-sales/index.html)
