Witdem
EN / DE
Back

Alongside

Witdem and LangSmith

Different jobs on the same agentic stack — LangChain tracing and evals vs product goals. Witdem does not replace LangSmith.

What LangSmith is strong at

LangSmith is observability and evaluation for LangChain and LangGraph. Teams pick it when they already ship LCEL chains or graph runs. They want datasets, evals, and production traces in one place. They do not pick it as a product-job scorecard.

Checkable strengths that matter in practice:

  • Deep fit with LangChain and LangGraph ecosystems
  • Datasets and evaluation workflows for iterating on chains and graphs
  • Production tracing for LCEL and graph runs
  • Strong when the team already lives in LangSmith day to day

Those skills stay useful with or without an outcomes layer. They answer how a LangChain run behaved. They also show how it scores on your eval set.

What Witdem adds

Witdem asks a different question. Did this run meet the product goal you named? That means a goal next to the code. It means evidence fields on the same run id. When useful, it also joins spend to that job.

A clean LangSmith trace can still leave the product question open. A green dataset eval can too. Finished LCEL steps can be true. A solid regression score can be true. The answer can still be empty, ungrounded, or outside the contract you shipped. Witdem closes that gap. It does not ask LangSmith to act as product analytics.

When you use both

The path is simple. Keep eval datasets and chain or graph debugging in LangSmith. Use Witdem when you need a clear answer to “did the job succeed?” on the same run. Wire both to one execution. Do not treat them as rivals for one slot. Product goal scoring lives in Witdem. Dataset evals stay where they already work.

Outcomes — product goal, evidence, spend×job (often Witdem) Tracing / evals — LCEL, graphs, datasets (often LangSmith)

The diagram is a map of jobs, not a ranking. Other tools can sit on either band. The labels mark where each product is typically strong.

Limits — when one alone is not enough

Witdem alone is not enough when you iterate on a LangGraph path. It is not enough for a bad LCEL step or dataset evals. You still need LangSmith traces and eval workflows. If debugging or regression evals are the main job this week, start with LangSmith.

LangSmith alone is not enough when product needs a named job outcome. You still need evidence and spend×job on one identity. Traces and dataset scores show how the run and the eval set behaved. They do not, by themselves, say whether the shipped product goal held. Use both when both questions matter.

What this means in practice

Picture one LangGraph or LCEL run from start to finish. You look at the same run twice. Once for traces and dataset evals. Once for the product job you promised.

First you open LangSmith. Steps look healthy. The eval set looks fine. That only says the chain or graph path completed and scored well on that set.

Then you ask the Witdem question. Was the named goal met? Do the evidence fields support pass or fail? A clean trace with an empty answer is still a product fail.

Keep both on one run id. Debug and regress in LangSmith. Score the job in Witdem. That is complementary use, not a swap.

FAQ

Does Witdem replace LangSmith? No. Witdem does not replace LangSmith. Keep traces and dataset evals. Add outcomes for the product job on the same run.

When do I use both? Use both when you need LangChain debug detail and a clear job result. Wire them to one execution identity.

What does Witdem add? A named product goal, evidence you can inspect, and spend joined to that job when useful.

Are dataset evals enough? No. Dataset scores show quality vs criteria. They do not, by themselves, say the shipped goal held.

Related

Tracing, evaluation, and product analytics · Witdem and Langfuse · Witdem and Arize Phoenix · Agentic workflow observability

Framework guides: Haystack · LangGraph · LangChain · OpenAI Agents

Get started Read the docs Guide
Witdem

Analytics for AI agents and multi-step AI applications.

Home Guide Haystack LangGraph LangChain OpenAI Agents Privacy Legal notice License