---
title: "LangChain and LangGraph explained: the essential toolkit | Runtime lines"
description: "The essential toolkit for Python AI developers: runnables and LCEL chains, chat models and tools, retrieval, and LangGraph agents with state, checkpoints and human-in-the-loop."
url: https://superintelligence.ro/langchain
alternate_en: https://superintelligence.ro/langchain.md
alternate_ro: https://superintelligence.ro/ro/langchain.md
---

# LangChain and LangGraph, the essential toolkit

LangChain gives you standard parts for LLM apps: chat models, messages, prompts, tools and retrievers, which all snap together as Runnables. LangGraph runs them as a graph with state, loops, memory and human approval. Step through four short programs, then read the parts, the patterns and the questions interviewers ask.

**Every program on this page really runs, without an API key.** The model is a fake chat model from `langchain_core.language_models` that replays scripted replies, so every printed output is real. In your app you write `model = init_chat_model("provider:model-name")` instead (with the provider’s package installed, such as `langchain-openai` or `langchain-anthropic`), and nothing else in the code changes.

## The building blocks

Everything in LangChain implements one interface, `Runnable`. Learn its methods once and you can call, compose, stream and test every part the same way: prompts, models, parsers, retrievers, tools, whole chains and compiled graphs.

### Chat models and messages

`init_chat_model("provider:model-name")` returns a chat model for any supported provider with the same API. A model takes a list of messages and returns an `AIMessage`.

- **SystemMessage**: instructions for the model.
- **HumanMessage**: what the user says.
- **AIMessage**: the reply. It can carry `tool_calls` and `usage_metadata` (token counts).
- **ToolMessage**: a tool’s result, linked to its call by `tool_call_id`.

In 1.x, `message.text` is a property that gives the plain text, and `message.content_blocks` gives typed blocks (text, reasoning, tool calls, images) in the same shape for every provider. Streaming yields `AIMessageChunk`s, which you can add together with `+`.

### Prompt templates

`ChatPromptTemplate` turns a dict into a list of messages. `MessagesPlaceholder` inserts a whole list, usually the chat history.
from langchain_core.messages import AIMessage, HumanMessage from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder prompt = ChatPromptTemplate.from_messages( [ ("system", "You are a {role}."), MessagesPlaceholder("history"), ("human", "{question}"), ] ) value = prompt.invoke( { "role": "tutor", "history": [HumanMessage("Hi"), AIMessage("Hello!")], "question": "What is LCEL?", } ) print([(m.type, m.content) for m in value.to_messages()]) [('system', 'You are a tutor.'), ('human', 'Hi'), ('ai', 'Hello!'), ('human', 'What is LCEL?')]

### Structured output and parsers

`model.with_structured_output(Schema)` returns a Runnable that gives you a validated object instead of a message. It uses the provider’s native JSON mode or tool calling under the hood. Pass `include_raw=True` to also get the raw message and any parsing error.
from langchain_core.language_models import GenericFakeChatModel from langchain_core.messages import AIMessage from pydantic import BaseModel class FakeChatModel(GenericFakeChatModel): def bind_tools(self, tools, **kwargs): return self class Movie(BaseModel): """A movie mentioned in the text.""" title: str year: int call = {"name": "Movie", "args": {"title": "Inception", "year": 2010}, "id": "c1"} model = FakeChatModel(messages=iter([AIMessage("", tool_calls=[call])])) extractor = model.with_structured_output(Movie) movie = extractor.invoke("Nolan's 2010 film about dreams") print(repr(movie)) Movie(title='Inception', year=2010)

Output parsers do the same job inside a chain: `StrOutputParser` for text, `JsonOutputParser` and `PydanticOutputParser` for models without native structured output.

### Tools

`@tool` turns a typed function into a tool: the name comes from the function, the description from the docstring and the argument schema from the type hints (or a Pydantic `args_schema`). The model sees only those three, so write them for the model.

- `model.bind_tools([...])` lets the model decide to call them.
- `tool.invoke(tool_call)` returns a `ToolMessage` with the right `tool_call_id`.
- `ToolNode` runs all the calls in the last `AIMessage` in parallel. By default it sends invalid-argument errors back to the model and re-raises other exceptions; `handle_tool_errors=True` returns every error as a message.

### RAG: documents, splitters, embeddings, retrievers

- **Loaders** read sources into `Document`s (`page_content` plus `metadata`).
- **Text splitters** (package `langchain-text-splitters`) cut them into chunks. `RecursiveCharacterTextSplitter` tries paragraphs, then lines, then words; `chunk_size` caps each chunk and `chunk_overlap` repeats a little text between neighbours so an idea isn’t cut in half.
- **Embeddings** turn text into vectors (`init_embeddings("provider:model-name")`); a **vector store** keeps them and finds the nearest ones.
- `store.as_retriever()` gives a **retriever**: a Runnable from a query string to a list of documents, so it drops straight into a chain.
from langchain_core.documents import Document from langchain_core.embeddings import DeterministicFakeEmbedding from langchain_core.language_models import FakeListChatModel from langchain_core.output_parsers import StrOutputParser from langchain_core.prompts import ChatPromptTemplate from langchain_core.runnables import RunnablePassthrough from langchain_core.vectorstores import InMemoryVectorStore docs = [ Document("LangGraph adds state, cycles and checkpoints.", metadata={"src": "a"}), Document("LCEL joins Runnables with the | operator.", metadata={"src": "b"}), ] store = InMemoryVectorStore.from_documents(docs, DeterministicFakeEmbedding(size=64)) retriever = store.as_retriever(search_kwargs={"k": 1}) prompt = ChatPromptTemplate.from_template( "Answer from the context only.\nContext: {context}\nQuestion: {question}" ) model = FakeListChatModel(responses=["(the model's grounded answer)"]) def format_docs(found): return "\n\n".join(d.page_content for d in found) rag = ( {"context": retriever | format_docs, "question": RunnablePassthrough()} | prompt | model | StrOutputParser() ) print(rag.invoke("What does LangGraph add?"))

The fake embeddings here are random vectors, so the ranking means nothing; swap in a real embedding model and the code stays the same.

## Chain, create_agent, or a custom graph?

Pick the simplest shape that fits. Each shape builds on the one before it, and all three are Runnables, so you can move up later without rewriting your tools or prompts.

-  Are the steps fixed, with no loop? **A chain (LCEL)**

Prompt, model, parser, maybe a retriever. Predictable, cheap and easy to test. Most extraction, classification, summarising and RAG apps are chains.
prompt | model | parser

-  Does the model choose tools until it’s done? **create_agent**

The standard tool-calling loop, built on LangGraph. Add behaviour with **middleware** instead of rewriting the loop: human approval, summarising long histories, retries, fallback models, call limits, PII redaction.
create_agent(model, tools, middleware=[...])

-  Do you need your own control flow? **A custom StateGraph**

Several agents, branches, parallel fan-out, approval at a specific step, long-running jobs that must survive restarts. You design the state, the nodes and the edges.
StateGraph(State).add_node(...)

### create_agent builds the agent loop for you

The same fake model and tool as in the stepper. The result is a compiled LangGraph graph with nodes `model` and `tools`. The `system_prompt` is sent with every model call but is not stored in the state.
from langchain.agents import create_agent from langchain_core.language_models import GenericFakeChatModel from langchain_core.messages import AIMessage, HumanMessage from langchain_core.tools import tool class FakeChatModel(GenericFakeChatModel): def bind_tools(self, tools, **kwargs): return self @tool def get_weather(city: str) -> str: """Get the current weather for a city.""" return f"Sunny, 21 C in {city}" call = {"name": "get_weather", "args": {"city": "Paris"}, "id": "call_1"} model = FakeChatModel( messages=iter([AIMessage("", tool_calls=[call]), AIMessage("Sunny, 21 C.")]) ) agent = create_agent(model, tools=[get_weather], system_prompt="You report weather.") result = agent.invoke({"messages": [HumanMessage("Weather in Paris?")]}) print(list(agent.get_graph().nodes)) print([m.type for m in result["messages"]]) ['__start__', 'model', 'tools', '__end__'] ['human', 'ai', 'tool', 'ai']

### Middleware: hooks around the loop

Middleware wraps the model and tool calls inside `create_agent`. Built-in ones, all in `langchain.agents.middleware`:

- `HumanInTheLoopMiddleware`: pause before chosen tools; a person can approve, edit, reject or respond.
- `SummarizationMiddleware`: summarise old messages when the history gets long.
- `ModelRetryMiddleware`, `ToolRetryMiddleware`, `ModelFallbackMiddleware`: retry with backoff, then switch models.
- `ModelCallLimitMiddleware`, `ToolCallLimitMiddleware`: cap calls per run or per thread.
- `PIIMiddleware`: redact or block personal data.

Write your own with the decorators `@before_model`, `@after_model`, `@wrap_model_call`, `@wrap_tool_call` and `@dynamic_prompt`.

### And when you need neither

One model call with no tools, no retrieval and no steps? The provider’s SDK, or a single `init_chat_model(...).invoke(...)`, is enough. Frameworks pay off when you swap providers, compose steps, stream, trace, or run loops with state. Don’t add layers you can’t explain.

## LangGraph core concepts

A LangGraph app is a **state** (a TypedDict, dataclass or Pydantic model), **nodes** that read it and return updates, and **edges** that decide which node runs next. It runs in **super-steps**: all nodes scheduled for a step run (in parallel if there are several), their updates are merged, then the next step starts.

### State and reducers

Nodes return **partial updates**, never the whole state. A key without a reducer is overwritten; a key annotated with a reducer is merged. `add_messages` appends new messages and replaces one whose `id` matches, which is how you edit a message.
import operator from typing import Annotated, TypedDict from langchain_core.messages import AIMessage, HumanMessage from langgraph.graph import add_messages class State(TypedDict): question: str # no reducer: a new value overwrites the old one log: Annotated[list[str], operator.add] # reducer: lists are concatenated messages: Annotated[list, add_messages] # append, or replace by message id print(operator.add(["start"], ["agent"])) old = [HumanMessage("Hi", id="1"), AIMessage("Draft", id="2")] print([m.content for m in add_messages(old, [AIMessage("Final", id="2")])]) print([m.content for m in add_messages(old, [AIMessage("More")])]) ['start', 'agent'] ['Hi', 'Final'] ['Hi', 'Draft', 'More']

`MessagesState` is the ready-made state with just `messages` and `add_messages`; subclass it to add keys.

### Nodes, edges and conditional edges

- A **node** is a function (sync or async) that takes the state and returns a dict of updates.
- `add_edge(a, b)` always goes from `a` to `b`. `START` and `END` mark the entry and exit.
- `add_conditional_edges(a, router)` calls `router(state)` and goes where it returns. Pass a list or dict of targets so the graph can be drawn.
- A node can also return `Command(goto="b", update={...})` to update state and route in one place.
- An edge back to an earlier node makes a **cycle**. That loop is the difference between a graph and a chain.

### Cycles and the recursion limit

Every super-step counts towards `recursion_limit`. When a run hits it, LangGraph raises `GraphRecursionError` instead of looping forever. Set it per call in the config.
from typing import TypedDict from langgraph.errors import GraphRecursionError from langgraph.graph import START, StateGraph class State(TypedDict): n: int builder = StateGraph(State) builder.add_node("loop", lambda s: {"n": s["n"] + 1}) builder.add_edge(START, "loop") builder.add_edge("loop", "loop") # a cycle with no exit graph = builder.compile() try: graph.invoke({"n": 0}, {"recursion_limit": 5}) except GraphRecursionError as e: print(type(e).__name__, str(e).splitlines()[0]) GraphRecursionError Recursion limit of 5 reached without hitting a stop condition. You can increase the limit by setting the `recursion_limit` config key.

Older releases defaulted to 25 steps. In langgraph 1.2 the default is 10,007 (`LANGGRAPH_DEFAULT_RECURSION_LIMIT`), and `create_agent` sets 9,999, so set your own limit for agents that could loop.

### Checkpointers, threads and time travel

Compile with a checkpointer and every super-step is saved as a **checkpoint** under the `thread_id` in the config. That gives you conversation memory, resumable runs after a crash, interrupts, and time travel: go back to an old checkpoint, change it, and run forward on a new branch.
from typing import TypedDict from langgraph.checkpoint.memory import InMemorySaver from langgraph.graph import START, StateGraph class State(TypedDict): n: int builder = StateGraph(State) builder.add_node("double", lambda s: {"n": s["n"] * 2}) builder.add_node("inc", lambda s: {"n": s["n"] + 1}) builder.add_edge(START, "double") builder.add_edge("double", "inc") graph = builder.compile(checkpointer=InMemorySaver()) config = {"configurable": {"thread_id": "t"}} print(graph.invoke({"n": 5}, config)) history = list(graph.get_state_history(config)) # newest first before_inc = next(c for c in history if c.next == ("inc",)) print(before_inc.values) fork = graph.update_state(before_inc.config, {"n": 100}) print(graph.invoke(None, fork)) {'n': 11} {'n': 10} {'n': 101}

Use `InMemorySaver` in tests and a database saver in production, such as `PostgresSaver` from `langgraph-checkpoint-postgres`.

### Interrupts

- `interrupt(value)` inside a node pauses the run and returns `value` to the caller under `__interrupt__`. It needs a checkpointer.
- Resume with `graph.invoke(Command(resume=answer), config)` on the same thread. The node **restarts from its first line**, and `interrupt()` returns `answer`.
- Because of the restart, keep side effects after the `interrupt()` call, or make them idempotent.
- `interrupt_before` and `interrupt_after` at compile time pause around whole nodes; they are mostly for debugging.

### Send: map-reduce and parallel branches

When the number of branches is only known at run time, return a list of `Send(node, input)` from a conditional edge. Each one runs the node with its own input, in parallel in the same super-step, and a reducer collects the results.
import operator from typing import Annotated, TypedDict from langgraph.graph import START, StateGraph from langgraph.types import Send class State(TypedDict): topics: list[str] summaries: Annotated[list[str], operator.add] def fan_out(state: State): return [Send("summarise", {"topic": t}) for t in state["topics"]] def summarise(item: dict): return {"summaries": [f"summary of {item['topic']}"]} builder = StateGraph(State) builder.add_node("summarise", summarise) builder.add_conditional_edges(START, fan_out, ["summarise"]) graph = builder.compile() print(graph.invoke({"topics": ["cats", "dogs", "owls"]})) {'topics': ['cats', 'dogs', 'owls'], 'summaries': ['summary of cats', 'summary of dogs', 'summary of owls']}

For a fixed set of branches you don’t need `Send`: add several edges out of one node and the targets run in parallel.

### Subgraphs

A compiled graph is a Runnable, so it can be a node in another graph. If both share state keys, pass it straight to `add_node`. If the schemas differ, call it inside a node function and map the state in and out. Subgraphs are how you build multi-agent systems from smaller, testable agents.
from typing import TypedDict from langgraph.graph import START, StateGraph class State(TypedDict): text: str def clean(state: State): return {"text": state["text"].strip()} inner = StateGraph(State) inner.add_node("clean", clean) inner.add_edge(START, "clean") cleaner = inner.compile() outer = StateGraph(State) outer.add_node("cleaner", cleaner) # a compiled graph is a node outer.add_node("shout", lambda s: {"text": s["text"].upper()}) outer.add_edge(START, "cleaner") outer.add_edge("cleaner", "shout") print(outer.compile().invoke({"text": " hi "})) {'text': 'HI'}

### Long-term memory: the Store

A checkpointer remembers one **thread**. A **Store** keeps JSON documents under namespaces, shared across threads: user preferences, learned facts. Nodes reach it, and the per-run `context`, through the `Runtime` argument.
from dataclasses import dataclass from langgraph.graph import START, MessagesState, StateGraph from langgraph.runtime import Runtime from langgraph.store.memory import InMemoryStore @dataclass class Context: user_id: str def remember(state: MessagesState, runtime: Runtime[Context]): ns = ("users", runtime.context.user_id) runtime.store.put(ns, "prefs", {"language": "Romanian"}) return {} builder = StateGraph(MessagesState, context_schema=Context) builder.add_node("remember", remember) builder.add_edge(START, "remember") store = InMemoryStore() graph = builder.compile(store=store) graph.invoke({"messages": []}, context=Context(user_id="ana")) print(store.get(("users", "ana"), "prefs").value) {'language': 'Romanian'}

With an embedding index configured, `store.search(namespace, query=...)` finds memories by meaning.

## Production essentials

The parts interviewers ask about once your demo works: seeing what happened, surviving flaky providers, stopping runaway loops, and testing without paying for tokens.

### LangSmith tracing

Set `LANGSMITH_TRACING=true` and `LANGSMITH_API_KEY` (and optionally `LANGSMITH_PROJECT`) and every chain, agent and graph run is recorded as a tree: each step’s inputs, outputs, latency, token usage and errors. No code changes. Wrap your own functions with `@traceable` from the `langsmith` package to include them.

### LangSmith evaluation

Keep a **dataset** of example inputs and reference outputs, run your app over it with `evaluate(target, data=..., evaluators=[...])`, and score each result with code checks or an LLM judge. Each run is an **experiment**, so you can compare prompts and models side by side before you ship, then run online evaluators on production traces.

### Retries, fallbacks and rate limits

`with_retry` retries a Runnable with exponential backoff; `with_fallbacks` tries the next Runnable when it still fails. Here the primary always times out:
from langchain_core.language_models import FakeListChatModel from langchain_core.runnables import RunnableLambda attempts = [] def flaky_model(prompt): attempts.append(prompt) raise TimeoutError("provider timed out") primary = RunnableLambda(flaky_model) backup = FakeListChatModel(responses=["Answer from the backup model"]) model = primary.with_retry(stop_after_attempt=3).with_fallbacks([backup]) print(model.invoke("hi").content) print(len(attempts)) Answer from the backup model 3

To stay under a provider’s rate limit, pass `rate_limiter=InMemoryRateLimiter(requests_per_second=...)` to the chat model. In agents, use the retry and fallback middleware.

### Limits and long conversations

- Set `recursion_limit`, and in agents `ModelCallLimitMiddleware` or `ToolCallLimitMiddleware`, so a confused model can’t loop and spend forever.
- Histories outgrow the context window: trim them with `trim_messages`, or summarise them with `SummarizationMiddleware`.
- Use a database checkpointer and choose `durability` (`"sync"`, `"async"` or `"exit"`) to trade safety for speed.

### Testing with fake models

`FakeListChatModel` replays strings, `GenericFakeChatModel` and `FakeMessagesListChatModel` replay whole `AIMessage`s, including `tool_calls`. Script the model, run the real chain or graph, and assert on the messages, the route taken and the final state, exactly like every program on this page. Unit tests stay fast, free and deterministic; check answer quality separately with LangSmith evaluations.

### LangChain, LangGraph, LangSmith

- **LangChain**: the parts (models, messages, tools, retrievers), LCEL, and `create_agent`.
- **LangGraph**: the runtime for stateful, looping, durable workflows. `create_agent` runs on it.
- **LangSmith**: tracing, evaluation and monitoring. It works with or without the other two.

## What to remember

**Everything is a Runnable**`invoke`, `batch`, `stream` and their async twins work on every part. `|` builds a `RunnableSequence`; a dict becomes a `RunnableParallel`.
**Messages are the interface**System, Human, AI and Tool messages. An `AIMessage` can ask for tools; each `ToolMessage` answers one call by `tool_call_id`.
**The model never runs tools**It returns `tool_calls`. Your code, `ToolNode` or `create_agent` runs them and sends the results back.
**An agent is a loop**Model, tools, model, until there are no tool calls. `create_agent` builds it; `StateGraph` lets you shape it.
**State plus reducers**Nodes return partial updates and reducers merge them. `add_messages` appends, or replaces by id.
**Checkpointer plus thread_id**That pair gives you memory, interrupts with `Command(resume=...)`, crash recovery and time travel. A Store remembers across threads.

## Interview questions: LangChain and LangGraph

Short answers you can say out loud. Try answering each question yourself before you open it.

### What is LangChain for, and when would you not use it?

LangChain gives you standard, provider-independent parts for LLM apps: chat models, messages, prompts, tools, retrievers and output parsing, all composable as Runnables, plus `create_agent` for tool-calling agents. It pays off when you compose steps, swap providers, stream, trace or build agents. For a single model call with no tools or retrieval, the provider’s SDK is simpler, and you shouldn’t add a layer you can’t explain.

### What is the difference between LangChain, LangGraph and LangSmith?

LangChain is the set of building blocks and the high-level `create_agent`. LangGraph is the low-level runtime for stateful workflows: graphs with state, cycles, checkpoints, interrupts and streaming; `create_agent` runs on it. LangSmith is the observability and evaluation platform, and it works with or without the other two.

### What is a Runnable, and what is LCEL?

A Runnable is anything with the standard interface: `invoke`, `batch`, `stream` and their async versions. Prompts, models, parsers, retrievers, tools and compiled graphs are all Runnables. LCEL, the LangChain Expression Language, composes them with `|`, which builds a `RunnableSequence` where each step’s output is the next step’s input. The composed chain is itself a Runnable, so it gets streaming, batching, async and tracing for free.

### What is the difference between `invoke`, `batch`, `stream` and `astream`?

`invoke` takes one input and returns one output. `batch` takes a list and runs the inputs concurrently on a thread pool, limited by `max_concurrency`. `stream` yields the output in chunks, such as tokens, as soon as they are ready. `astream` and the other `a` methods are the async versions for asyncio servers, so a slow model call doesn’t block other requests.

### What do `RunnableParallel`, `RunnablePassthrough` and `RunnableLambda` do?

`RunnableParallel` runs several Runnables on the same input and returns a dict of their results; a plain dict inside a chain is turned into one automatically. `RunnablePassthrough` passes its input through unchanged, and `RunnablePassthrough.assign` adds computed keys to an input dict. `RunnableLambda` wraps a plain Python function so it can sit in a chain. The classic RAG chain uses all three ideas: `{"context": retriever | format, "question": RunnablePassthrough()} | prompt | model`.

### What are the message types, and what does each one mean?

`SystemMessage` carries instructions, `HumanMessage` the user’s input, and `AIMessage` the model’s reply, which can include `tool_calls` and `usage_metadata`. `ToolMessage` carries a tool’s result and must have the `tool_call_id` of the call it answers. A chat model takes a list of messages and returns an `AIMessage`; in 1.x, `.text` gives the text and `.content_blocks` gives typed blocks in the same shape for every provider.

### How does tool calling work, step by step?

You bind tools with `model.bind_tools(tools)`, which sends each tool’s name, description and JSON schema with the request. The model doesn’t run anything: it returns an `AIMessage` with `tool_calls`, each with a name, args and an id. Your code runs each tool and appends a `ToolMessage` with the matching `tool_call_id`, then calls the model again with the whole history. When the model replies without tool calls, you’re done.

### `bind_tools` or `with_structured_output`: when do you use each?

Use `bind_tools` when the model should decide whether to act and which tools to call, as in an agent. Use `with_structured_output(Schema)` when you always want data in a fixed shape, such as extraction or classification: it forces the format with the provider’s JSON mode or a tool call, then parses the result into your Pydantic model, TypedDict or dict. `include_raw=True` also returns the raw message and any parsing error.

### How does an agent loop work?

The model gets the conversation and the tool schemas. If it returns tool calls, the tools run and their results are appended as `ToolMessage`s, then the model is called again. The loop ends when the model answers without tool calls, or when a limit stops it. In LangGraph that’s two nodes, model and tools, a conditional edge such as `tools_condition`, and an edge from tools back to the model.

### When would you use `create_agent`, and when a hand-built LangGraph graph?

`create_agent(model, tools, ...)` gives you the standard tool-calling loop on LangGraph, with checkpointing, streaming and middleware for things like human approval, summarisation, retries and call limits. Use it whenever the model choosing tools in a loop is the right shape. Build your own `StateGraph` when you need custom control flow: fixed steps mixed with agent steps, several agents, parallel fan-out, or approval at one specific step.

### What is middleware in `create_agent`?

Hooks that run around the agent’s model and tool calls, so you change behaviour without rewriting the loop. Built-in ones include `HumanInTheLoopMiddleware`, `SummarizationMiddleware`, `ModelRetryMiddleware`, `ModelFallbackMiddleware`, `ModelCallLimitMiddleware`, `ToolCallLimitMiddleware` and `PIIMiddleware`. You write your own with decorators such as `@before_model`, `@after_model`, `@wrap_model_call` and `@dynamic_prompt`.

### What are `StateGraph`, state and reducers? Why `add_messages`?

A `StateGraph` is built from a state schema, usually a TypedDict. Nodes return partial updates, and each key’s **reducer** decides how an update is merged: no reducer means overwrite, `operator.add` means concatenate. `add_messages` appends new messages and replaces any with the same id, so nodes can return just the new message and parallel nodes don’t overwrite each other. `MessagesState` is the ready-made state with a `messages` key using it.

### How do conditional edges and cycles work?

`add_conditional_edges(node, router)` calls `router(state)` after the node and goes to the node it returns, or to `END`. An edge back to an earlier node creates a cycle, which is how agents loop. A node can also return `Command(goto=..., update=...)` to route and update in one step. Every loop needs an exit condition, and the recursion limit is the safety net.

### What is the recursion limit, and what happens when you hit it?

It caps the number of super-steps in one run. When it’s reached, LangGraph raises `GraphRecursionError` instead of looping forever. Set it per call with `config={"recursion_limit": n}`. Older versions defaulted to 25; langgraph 1.2 defaults to 10,007 and `create_agent` uses 9,999, so set your own for agents, or add call-limit middleware.

### What does a checkpointer do, and what is a thread?

A checkpointer saves the graph’s state after every super-step. A thread is one conversation or job, identified by `thread_id` in `config["configurable"]`; all its checkpoints are stored under that id. Invoke again with the same thread and the graph continues from the saved state, which gives you conversation memory, recovery after a crash, interrupts and time travel. Use `InMemorySaver` in tests and a Postgres or SQLite saver in production.

### What is time travel in LangGraph?

Because every step is checkpointed, you can list past states with `get_state_history(config)`. Invoking with an older checkpoint’s config replays from that point. `update_state` on an old checkpoint creates a fork you can run forward with different values. It’s used for debugging, for retrying from a bad step, and for letting users edit an earlier turn.

### How do `interrupt` and `Command(resume=...)` work, and what’s the catch?

Calling `interrupt(value)` in a node saves the state and stops the run; the caller gets `value` under `__interrupt__`. To continue, invoke the same thread with `Command(resume=answer)`. The catch: the node restarts from its first line and `interrupt()` then returns `answer`, so any code before it runs twice. Keep side effects after the interrupt or make them idempotent. Interrupts need a checkpointer.

### Short-term or long-term memory: checkpointer or Store?

Short-term memory is the state of one thread, kept by the checkpointer: the messages of this conversation. Long-term memory has to outlive threads, such as user preferences or facts learned earlier, so it goes in a **Store**: JSON documents under namespaces like `("users", user_id)`, optionally searchable by meaning. Nodes reach the store through the `Runtime` argument.

### What is `Send` for, and how do parallel branches work?

Nodes scheduled in the same super-step run in parallel, so several edges out of one node already give you fixed parallel branches. When the number of branches is only known at run time, a conditional edge returns a list of `Send(node, input)`, and each one runs the node with its own input, the map step of map-reduce. A reducer such as `operator.add` on the result key collects their outputs.

### What are subgraphs, and why use them?

A compiled graph is a Runnable, so it can be a node inside another graph. If the parent and child share state keys, you add it directly; if not, you call it inside a node and map the state in and out. Subgraphs let you build and test agents separately and then combine them, which is the usual way to build multi-agent systems. A node in a subgraph can route in the parent with `Command(graph=Command.PARENT, goto=...)`.

### What streaming modes does LangGraph have?

`values` streams the full state after each step, and `updates` streams only what each node returned. `messages` streams LLM tokens with metadata saying which node produced them, and `custom` streams whatever nodes send with `get_stream_writer()`. There are also `checkpoints`, `tasks` and `debug`; pass a list to get several as `(mode, data)` tuples.

### How do you stream an agent’s answer to a web UI?

Use `astream` with `stream_mode="messages"` to get tokens as they’re generated, often combined with `"updates"` to show progress such as which tool is running. Forward the chunks to the browser over server-sent events or a WebSocket. Token streaming works even if the node calls `invoke`, because LangGraph listens to the model’s callbacks.

### How do you build RAG with LangChain?

Indexing: load sources into `Document`s, split them into chunks with a text splitter, embed the chunks and store them in a vector store. Answering: a retriever finds the most similar chunks for the question, and a chain puts them into the prompt and asks the model to answer from that context only. For harder questions you make retrieval a tool, so an agent can search several times and rewrite its query.

### How do you choose chunk size and overlap?

Chunks should be big enough to hold one complete idea and small enough that a retrieved chunk is mostly relevant; a few hundred to a thousand or so tokens is a common start. Overlap, often 10 to 20 percent, repeats text across neighbouring chunks so an idea cut at a boundary is still found. `RecursiveCharacterTextSplitter` splits on paragraphs, then lines, then words, to keep chunks natural. Then measure retrieval quality on real questions and tune.

### What is a retriever, and how is it different from a vector store?

A vector store stores embeddings and runs similarity search. A retriever is a Runnable interface: a query string in, a list of `Document`s out. `vector_store.as_retriever(search_kwargs={"k": 4})` wraps a store, but a retriever can also be keyword search, a web search or a hybrid of several. Because it’s a Runnable, it drops into any chain.

### What is LangSmith used for?

Tracing: with `LANGSMITH_TRACING=true` and an API key, every run is recorded as a tree of steps with inputs, outputs, latency, tokens and errors, so you can see exactly what the model was sent. Evaluation: you keep datasets of examples, run your app over them with code or LLM-judge evaluators, and compare experiments before shipping. It also monitors production traffic and can run online evaluations on it.

### How do you handle retries, fallbacks and rate limits?

`runnable.with_retry(stop_after_attempt=3)` retries with exponential backoff, and `with_fallbacks([backup])` tries another model or chain when it still fails. Pass `rate_limiter=InMemoryRateLimiter(requests_per_second=...)` to a chat model to stay under provider limits. In agents, `ModelRetryMiddleware`, `ToolRetryMiddleware` and `ModelFallbackMiddleware` do the same jobs.

### How do you test LangChain and LangGraph code without calling a real model?

Use the fake chat models in `langchain_core`: `FakeListChatModel` replays strings, and `GenericFakeChatModel` or `FakeMessagesListChatModel` replay whole `AIMessage`s, including tool calls. Run the real chain or graph with an `InMemorySaver` and assert on the messages, the route taken and the final state. That keeps unit tests fast and deterministic; you judge answer quality separately with LangSmith evaluations.

### How do you handle a conversation that no longer fits in the context window?

Trim it, keeping the system message and the most recent messages, with `trim_messages`; summarise older turns into a short message, which `SummarizationMiddleware` does for agents; or move lasting facts to a Store and retrieve them when needed. When trimming, never split an `AIMessage` with tool calls from its `ToolMessage`s, or the provider rejects the history.

### What happens when a tool raises an exception?

In a graph, `ToolNode` by default turns invalid-argument errors into a `ToolMessage` so the model can correct its call, and re-raises other exceptions. Set `handle_tool_errors=True`, or pass a function, to send every error back to the model as a message. In `create_agent`, retry middleware can retry flaky tools first.

### What does `init_chat_model` do?

It creates a chat model from a string like `"provider:model-name"`, loading the right integration package, such as `langchain-openai` or `langchain-anthropic`. Every provider then has the same interface, so you can switch models through configuration without changing code. `create_agent` also accepts the same string directly.

### How do multi-agent systems work in LangGraph?

Each agent is a node or a subgraph with its own prompt and tools. In a supervisor design, one agent routes work to the others, often by calling them as tools; in a handoff design, agents pass control directly with `Command(goto=...)`. Shared state or messages carry the context between them. Start with one agent and add more only when tools or instructions clearly separate.

Every output on this page was printed by the programs in `verify/langchain/`, run with langchain 1.4.3, langchain-core 1.6.7 and langgraph 1.2.14 on Python 3.14. The models are the fake chat models from `langchain_core.language_models`; message ids are left out of the drawings.
