# How Candidly Built State-Aware Agent Harnesses with LangSmith

B. Levine,  
P. Hendershott  
June 29, 2026

13 min

_Candidly helps people make high-stakes financial decisions around debt payoff, savings, retirement, benefits, and education costs. Cait is our AI financial planner—a conversation agent that helps compare savings strategies, repayment options, understand deadlines, evaluate tradeoffs, and decide what to do next._

For example:
- "How do I balance saving for the future with paying off my debt?"
- "How can I create more space in my budget? Can I afford college for my child without sacrificing retirement?"
- "Should I pause my 401(k) contributions to pay off my student loans faster?"

## From _Ex-Post_ Evaluations to Live Steering

Most conversational assistants are judged after the fact, by how the conversation ended. _Did the user get an answer? Did they complete the task? Did they return, click through, or take the next step?_

To optimize for resolution during the conversation, the agent harness needs a turn-level view of where the interaction is and which response levers can move it forward. We build that view by reading the partial trace, inferring the user’s current state, and selecting response features based on how similar conversations shifted in the past.

This post builds on a [formal research paper](https://go.getcandidly.com/State-dependent-interventions-IO-HMM) and a multi-month analysis of production Cait conversations.

## Can We Predict How a Conversation Will End?

To assess conversation-level outcomes at scale, we built a hybrid labeling pipeline over production Cait traces and tracked the resulting labels in LangSmith. Deterministic rules handled clear cases like explicit frustration, no reply after Cait's first message, or product follow-through such as linking an account or starting a savings plan.

Ambiguous cases went to LLM-as-judge evaluators, routed by conversation pattern. Each verdict attaches to its thread as feedback. We calibrated the pipeline against a human-labeled LangSmith dataset, reaching 92.3% agreement.

Then we trained a model to predict that label from features computed directly off the trace. A few examples of these features, split between what Cait did and how the user responded:

**What Cait did**
- **Q/A alignment**: the lexical overlap between Cait's response and the user prompt that preceded it. High alignment is one of the strongest predictors of resolution in our data; and low alignment is a defining signature of the failing state.
- **Topic continuity**: the semantic persistence of Cait’s own responses. For Cait response at turn _t_, it measures the semantic similarity between Cait’s current response _aₜ_ and Cait’s previous response _aₜ₋₁._ Continuity and coherence across the trajectory strongly predicts resolution.

**How the user responded**
- **Message length**: Longer messages predict resolution; low or one-word replies predict a user on the way out.
- **Caps ratio**: the share of the user message in capital letters (a frustration signal that predicts abandonment).

A gradient-boosted model trained on these features separated resolved from abandoned conversations apart at **0.90 AUC** (0.5 is chance, 1.0 is perfect). Resolution versus abandonment was predictable from signals in the traces.

## Turning Traces into a State Model

We care about modeling the complex interaction between the user, the agent, and the conversation context. The prediction results told us that conversation outcomes are learnable from the trace. To make that signal useful while Cait is still talking to the user, we need a model of how conversations unfold, including: how the observable signals accumulate over turns, which patterns tend to recur, and which parts of Cait’s behavior are associated with movement toward or away from resolution.

A useful model has to do three things:

1. Represent the conversation as an ordered trajectory we can interpret and act on.
2. Separate user-side signals from agent-side features, because they play different roles. User behavior tells us what state the conversation is in, while agent behavior is the lever the system can move.
3. Learn the mapping from signals to states from the data itself, so the states reflect the patterns that actually appear, rather than categories we impose in advance.

A model with those properties gives us more than a retrospective score. It can read the partial trace, summarize where the conversation appears to be, and connect that readout to response features Cait controls. An Input-Output Hidden Markov Model meets all three requirements.

### Inside the IO-HMM

The IO-HMM separates each turn into two pieces:

- **User-side signals** are emissions. They are the observable behaviors we use to infer the user's current engagement state.
- **Agent-side features** are transition inputs. They are the controllable response characteristics that condition where the conversation moves next.

The model estimates where the conversation is likely to move next, given where it is now and how the agent just responded.

The model is fit across thousands of conversations using expectation-maximization.

The important design choice is the separation. User behavior is used to read the state of the conversation. Agent behavior is used to estimate transitions between states. That separation turns traces from a record of what happened into a model of what is driving the conversation.

## Four Engagement States, and Why Averages Hide Them

The model recovered four interpretable engagement states:

| State       | Share of turns | What it looks like                                                            | Outcome signature            |
|-------------|----------------|-----------------------------------------------------------------------------|-------------------------------|
| Engaged     | 53%            | Moderate, specific back-and-forth; strong alignment between the user's need and Cait's response | Highest resolution rate       |
| Detailed    | 7%             | Long, structured user messages with substantive financial context            | High resolution rate          |
| Guided      | 17%            | Shorter user replies; Cait carries more of the explanatory load             | High resolution rate; longest conversations |
| Disengaging | 23%            | Short user messages and weak alignment between the user's language and Cait's responses | Lowest resolution rate         |

The states differ in both behavior and outcomes. They map to whether users get their question resolved, continue exploring products, or abandon the conversation. Resolution ranges from about 78% in the most engaged state to about 30% in the disengaging state.

## Wiring the Policy into the Harness

A fitted model is not yet a policy. To act on it, the model has to run inside the harness, every change it drives has to be recorded, and its effect has to be verified before launch and measured after.

**State inference in the request path.** The model runs on every turn, fast enough to shape the next response. We freeze the fitted parameters, compute the user's current state from the signals already in the trace, and write that state back onto the trace as metadata, inside the latency budget of a single reply.

LangSmith closes the loop; traces become state-aware eval data, eval data becomes prompt-policy experiments, and production monitoring tells us whether the policy changed the trajectory.

## Building Better Agents: Evaluation as a Control Signal

Nothing about that is specific to financial conversations, or to conversations at all. The recipe requires three ingredients. 1.) an outcome that's only observed at the end, 2.) turn-level signals computable from the trace, and 3.) behaviors the agent controls. Any multi-turn agent with a goal has all three.

With that state estimate on every turn of every trace, evaluation becomes part of the operating loop: the trace records what happened, the model reads where the interaction is, the policy chooses the response, and the next trace shows whether that choice moved the conversation in the right direction.
