Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Multi-Agent Research Assistant with LangGraph

Authored by Dr. Tiziana Ligorio for AI Agents - CSCI 395.32 taught at Hunter College of The City University of New York
Adapted from: Large Language Model Agents, Jerin George Mathew & Jacopo Rossi, Springer 2025

In this tutorial, we build a research assistant that uses multiple agents to streamline the process of finding and filtering academic research papers. This demonstrates a multi-agent system using the LangGraph framework.

The system consists of four specialized agents:

  1. Search Agent — Queries arXiv to find academic papers matching the user’s query

  2. Filter Agent — Evaluates the relevance of retrieved papers and adds relevant ones to the filtered papers list

  3. Query Refinement Agent — Refines the search query to improve results when the current query yields insufficient relevant papers

  4. Supervisor Agent — Decides whether the workflow should finalize (enough relevant papers found) or continue refining the query

Workflow

Multiagent System - web

Stopping Criteria

The Supervisor Agent finalizes the workflow when at least 3 papers have been identified with a relevance score ≥ 0.7. Otherwise, the Query Refinement Agent generates an improved query and the search process iterates.

Installs and Imports

%%capture hides the output

OpenRouter is a unified API that provides access to various LLMs through a single interface. It offers a generous free tier and affordable token usage for minimal cost, making it ideal for learning and experimentation.

If you already pay for other LLM providers or prefer to use a different service, you are welcome to adapt the code accordingly.

Setup your API Key

Step 1 — Get an OpenRouter API key

For this demo we will use an LLM via OpenRouter, which requires an API key.

  1. Go to https://openrouter.ai

  2. Sign in (or create an account if you don’t have one)

  3. Once logged in, navigate to https://openrouter.ai/settings/keys

  4. Click Create Key

  5. Give the key a name, e.g. colab-multiagent_langgraph

  6. Copy the key immediately (you won’t be able to see it again)

Important: Treat this key like a password. Do not share it, paste it into notebooks, or commit it to GitHub.

Step 2 — Add a secret in Colab (UI)

  1. On the left sidebar, click 🔑 Secrets

  2. Add a new secret:

  • Name: OPENROUTER_API_KEY

  • Value: your actual API key

  1. Toggle the switch to the left to give notebook access (you should see a checkmark)

If running locally — Add a secret in .env

  1. Create a .env file in the project root:

touch .env

  1. Add the following (replace with your own key):

OPENROUTER_API_KEY=your_openrouter_key_here.

Important: Never paste API keys into code cells.

Load the API Key

In Colab:

Uncomment and run the cell below if you’re using Google Colab.

Locally:

True
OPENROUTER_API_KEY present: True

Define the State Schema

In a multi-agent system, the state serves as the shared memory through which agents communicate and coordinate. Each agent reads from and writes to this common structure, enabling them to build on each other’s work without direct interaction.

In LangGraph, the state is a shared data structure that flows through the graph and gets updated by each node. We define it as a TypedDict to specify what fields exist and their types.

When a node returns a dictionary, LangGraph merges it into the current state:

  • For regular fields, the returned value replaces the existing value

  • For fields using Annotated with a reducer (like operator.add), the returned value is combined with the existing value

This is important for our workflow:

  • papers gets replaced on each search (we only want the current iteration’s results)

  • filtered_papers accumulates across iterations (we want to keep all relevant papers found so far)

Define the Agents

Search Agent

The Search Agent is responsible for querying academic paper databases to find papers matching the user’s research query. It uses the arXiv API to search for papers and returns structured metadata for each result.

Input: Takes the current query from the state
Output: Returns a list of papers with metadata (title, authors, abstract, URL, publication date)
Tools: arXiv API client

The agent does not use an LLM — it’s a straightforward API call that retrieves papers based on keyword matching. The LLM-based reasoning happens in the Filter Agent, which evaluates relevance.

Search Agent: Found 10 papers for query 'multi-llm-agent reinforcement learning'
10
2024-09-27: ARLBench: Flexible and Efficient Benchmarking for Hyperparameter Optimization in Reinforcement Learning
2025-06-24: Causal-Paced Deep Reinforcement Learning
2018-07-13: Exploring Hierarchy-Aware Inverse Reinforcement Learning
2024-01-14: Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
2023-01-19: A Tutorial on Meta-Reinforcement Learning
2018-09-25: Anderson Acceleration for Reinforcement Learning
2024-06-07: Stabilizing Extreme Q-learning by Maclaurin Expansion
2019-09-26: MERL: Multi-Head Reinforcement Learning
2025-08-09: Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
2019-04-20: Compression and Localization in Reinforcement Learning for ATARI Games
10

Filter Agent

The Filter Agent evaluates the relevance of each paper retrieved by the Search Agent. Unlike the Search Agent, this agent uses an LLM to reason about semantic relevance — determining whether a paper’s content actually addresses the user’s research question, not just whether it contains matching keywords.

Input: Takes papers (raw search results) and query from the state
Output: Returns papers that score ≥ 0.7 relevance, each with a relevance_score field added
LLM: Uses gpt-4o-mini via OpenRouter for cost-effective reasoning

The agent prompts the LLM to return a JSON object with a relevance score (0.0–1.0) and justification for each paper. Only papers meeting the threshold are added to filtered_papers.

Filter Agent: 1/10 papers passed relevance threshold (>= 0.7)
1
'Small LLMs Are Weak Tool Learners: A Multi-LLM Agent'
10
ARLBench: Flexible and Efficient Benchmarking for Hyperparameter Optimization in Reinforcement Learning 0.4 The paper discusses hyperparameter optimization in reinforcement learning, which is related to the broader topic of reinforcement learning but does not specifically address multi-LLM-agent reinforcement learning.
Causal-Paced Deep Reinforcement Learning 0.4 The paper discusses reinforcement learning and curriculum learning, which are related to multi-agent reinforcement learning, but it does not specifically address multi-LLM agents or their integration.
Exploring Hierarchy-Aware Inverse Reinforcement Learning 0.4 The paper discusses inverse reinforcement learning and hierarchical strategies, which are related to reinforcement learning concepts, but it does not specifically address multi-LLM-agent systems.
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent 0.8 The paper discusses a multi-LLM agent framework that addresses tool learning, which is relevant to multi-LLM-agent reinforcement learning, particularly in the context of task planning and execution.
A Tutorial on Meta-Reinforcement Learning 0.4 The paper discusses meta-reinforcement learning, which is related to reinforcement learning but does not specifically address multi-LLM-agent systems.
Anderson Acceleration for Reinforcement Learning 0.4 The paper discusses reinforcement learning and introduces a method that could be applied to it, but it does not specifically address multi-LLM-agent reinforcement learning.
Stabilizing Extreme Q-learning by Maclaurin Expansion 0.4 The paper discusses reinforcement learning and introduces a method related to Q-learning, which is relevant to the broader topic of multi-agent reinforcement learning, but it does not specifically address multi-LLM agents or their integration.
MERL: Multi-Head Reinforcement Learning 0.4 The paper discusses reinforcement learning and introduces a framework (MERL) that could relate to multi-agent systems, but it does not specifically address multi-LLM-agent reinforcement learning, making it only somewhat relevant.
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code 0.6 The paper discusses multi-agent systems involving LLMs and their application in code generation, which relates to multi-LLM-agent reinforcement learning, but it does not explicitly focus on reinforcement learning aspects.
Compression and Localization in Reinforcement Learning for ATARI Games 0.3 The paper discusses reinforcement learning and model compression, but it does not specifically address multi-LLM-agent reinforcement learning, making it only tangentially relevant.

Supervisor Agent

The Supervisor Agent is the decision-maker that controls the workflow. After the Filter Agent evaluates papers, the Supervisor checks whether we have enough relevant results or need to refine the query and search again.

Input: Takes filtered_papers and iteration from the state
Output: Returns a decision field: either "end" or "refine"
No LLM required, this is pure conditional logic, not reasoning.

Decision Logic:

  1. If filtered_papers has ≥3 papers → "end" (success)

  2. If iteration ≥ 3 → "end" (max attempts reached, return what we have)

  3. Otherwise → "refine" (try again with a refined query)

Query Refinement Agent

The Query Refinement Agent improves the search query when the current results are insufficient. It uses an LLM to reason about why the previous query didn’t yield enough relevant papers and how to improve it.

Input: Takes query, all_evaluations, and iteration from the state
Output: Returns an updated query string and increments iteration
LLM: Uses gpt-4o-mini via OpenRouter to analyze feedback and generate better queries

The agent examines the evaluation feedback (why papers were rejected) and uses that insight to craft a more targeted query. For example, if many papers were rejected for being too theoretical, it might add terms like “applied” or “practical”.

Current state before refinement:
  query: 'multi-llm-agent reinforcement learning'
  iteration: 0
  all_evaluations: 10 items

Query Refinement Agent: 'multi-llm-agent reinforcement learning' → 'multi-llm-agent reinforcement learning framework tool learning task planning execution'

Refinement result:
  new query: 'multi-llm-agent reinforcement learning framework tool learning task planning execution'
  new iteration: 1

Build the Graph

Now that we have all four agents defined, we wire them together into a LangGraph StateGraph. The graph defines:

  1. Nodes — Each agent function becomes a node in the graph

  2. Edges — Define the flow between nodes (which agent runs after which)

  3. Conditional Edges — Allow dynamic routing based on state (the Supervisor’s decision)

Recall, the workflow follows this pattern:
Multiagent System - web

Graph compiled successfully!

Run the Workflow

Now we can run the complete workflow by invoking the compiled graph with an initial state. The graph will:

  1. Start with the search agent

  2. Filter results for relevance

  3. Check if we have enough papers (Supervisor)

  4. If not, refine the query and repeat

  5. Continue until we have ≥3 relevant papers or hit the max iteration limit

Starting workflow with query: 'multi-llm-agent reinforcement learning'
============================================================
Search Agent: Found 10 papers for query 'multi-llm-agent reinforcement learning'
Filter Agent: 1/10 papers passed relevance threshold (>= 0.7)
Supervisor Agent: REFINE - Refining query (iteration 1)
Query Refinement Agent: 'multi-llm-agent reinforcement learning' → 'multi-llm-agent reinforcement learning collaboration tool learning'
Search Agent: Found 10 papers for query 'multi-llm-agent reinforcement learning collaboration tool learning'
Filter Agent: 0/10 papers passed relevance threshold (>= 0.7)
Supervisor Agent: REFINE - Refining query (iteration 2)
Query Refinement Agent: 'multi-llm-agent reinforcement learning collaboration tool learning' → 'multi-agent reinforcement learning collaboration tools for large language models'
Search Agent: Found 10 papers for query 'multi-agent reinforcement learning collaboration tools for large language models'
Filter Agent: 3/10 papers passed relevance threshold (>= 0.7)
Supervisor Agent: END - Success: Found 4 relevant papers (>= 3 required)
============================================================

Workflow complete!
Result: Success: Found 4 relevant papers
Final query: 'multi-agent reinforcement learning collaboration tools for large language models'
Total iterations: 2
Relevant papers found: 4
Relevant Papers Found:
------------------------------------------------------------

1. Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
   Published: 2024-01-14
   Relevance: 0.8
   URL: http://arxiv.org/abs/2401.07324v3

2. Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
   Published: 2025-09-20
   Relevance: 0.7
   URL: http://arxiv.org/abs/2509.16679v1

3. Hierarchical Multi-agent Large Language Model Reasoning for Autonomous Functional Materials Discovery
   Published: 2025-12-15
   Relevance: 0.7
   URL: http://arxiv.org/abs/2512.13930v1

4. Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
   Published: 2024-12-06
   Relevance: 0.8
   URL: http://arxiv.org/abs/2412.05449v1

Some Suggestions for Improvements

  1. Duplicate papers — Papers may appear more than once in the final results. This happens because filtered_papers accumulates across iterations — if the same paper is found in multiple searches and passes the relevance threshold each time, it gets added again.

    Exercise for the reader: Modify the workflow to prevent duplicate papers. Consider deduplicating by paper URL or title in the filter_agent, keeping track of already-seen paper IDs in the state, or deduplicating at the end before displaying results.

  2. Multiple search sources — More search agents may be added to use different academic paper sources: Semantic Scholar, PubMed, OpenAlex, CrossRef, IEEE Xplore, ACM Digital Library, and others.

    Exercise for the reader: Implement additional agents for different searches (e.g., semantic_scholar_agent, pubmed_agent, etc.) and modify the graph to run multiple search agents in parallel. Consider how the state schema should change, whether the filter agent should treat sources differently, and how to handle cross-database duplicates. Should all sources always be searched? What logic determines which sources to use — the research domain, the query keywords, the iteration number? Or should an LLM agent decide?

  3. Error handling and retries — The workflow currently assumes API calls succeed. What happens if arXiv is down, rate limits us mid-workflow, or returns malformed data?

    Exercise for the reader: Add try/except blocks and graceful degradation so the workflow can recover from transient failures.

  4. Citation following — Once relevant papers are found, expanding the search to include papers they cite (references) or papers that cite them (citations) could surface important related work.

    Exercise for the reader: Implement a citation_agent that takes the filtered papers and queries a citation API (e.g., Semantic Scholar) to find connected papers. How should this agent integrate into the existing graph?

  5. Summarization agent — After finding relevant papers, a summarization step could help users quickly understand the landscape.

    Exercise for the reader: Add a final agent that synthesizes the findings: generate a research summary, identify common themes across papers, or produce a literature review outline.

  6. User feedback loop — The current workflow is fully automated. Allowing user input during execution could improve results.

    Exercise for the reader: Modify the workflow to pause and ask the user to mark papers as relevant or irrelevant. Use this feedback to adjust subsequent searches or filter criteria.

  7. Export functionality — Researchers need results in formats compatible with their tools.

    Exercise for the reader: Add an export step that outputs results to BibTeX for LaTeX, CSV for spreadsheets.

  8. Checkpointing — Long-running workflows may be interrupted. LangGraph supports persistence and checkpointing.

    Exercise for the reader: Configure the workflow to save state periodically so it can be resumed if interrupted. See the LangGraph documentation on persistence.