If your agent cannot search the live web, inspect real sources, and cite where its answer came from, it is usually guessing from stale model knowledge or from whatever documents you manually gave it.
For Day 1 of 30 Days of Search for AI Agents, we are going to build the simplest useful version of a search-powered agent:
A small Python agent that takes a question, searches the live web, and returns source-backed context that an LLM can use to answer with citations. No framework. No complicated orchestration. No giant RAG pipeline.
Just:
Quick summary
AI agents need live search when they need current, external, or source-grounded information. In this tutorial, you will build a small Python search agent using Valyu. The agent takes a user question, searches the live web, returns relevant source snippets, and formats them as context for an LLM to produce a cited answer.
To add live web search to an AI agent, give the agent a search tool that can retrieve fresh sources, preserve URLs, and pass source content into the model prompt. Valyu provides a search API that agents can use to query the live web and specialized knowledge sources, making answers more current, verifiable, and citation-ready.
Who this is for
This guide is for developers building:
- AI agents
- research assistants
- coding assistants
- financial research tools
- scientific literature agents
- market intelligence bots
- documentation assistants
- cited answer systems
If your agent needs to answer questions about the outside world, it needs search.
Why AI agents need search
An AI agent without search has three big problems.
1. It goes stale
The model may not know what happened yesterday, this morning, or five minutes ago.
2. It lacks source grounding
If an answer matters, the user needs to know where it came from.
3. It cannot inspect external sources on demand
Many real workflows need fresh pages, docs, filings, papers, datasets, or news.
For example, these are bad fits for static model memory:
- What did this company say in its latest 10-K?
- What are the newest papers on this biomedical topic?
- What changed in this API’s documentation?
- Which clinical trials are recruiting right now?
- What are the latest developments in inference data centers?
- What are people saying today about this open-source project?
These are search problems. And increasingly, they are agent search problems.
What we are building
Today we will build a tiny search agent that can answer questions like:
What are the latest AI inference datacenter projects?
The agent will:
- Accept a user question
- Search the live web with Valyu
- Return relevant source snippets
- Format those results as context for an LLM
- Preserve source URLs so the final answer can cite them
What is Valyu?
Valyu is a search & deepresearch API built for AI knowledge work.
It gives agents access to:
- live web search
- content extraction from URLs
- cited answers
- DeepResearch reports
- academic sources like arXiv, BiorXviv, chemrXviv, PubMed, etc
- finance sources like SEC filings and market data
- biomedical and life sciences sources
- patents
- clinical trials
- chemistry and drug discovery databases
- specialized datasets for knowledge work
For today, we will start with the simplest thing: live web search.
Install the SDK
First, install the Python SDK:
pip install valyu
Then get an API key from:
https://platform.valyu.ai/
Set it as an environment variable:
export VALYU_API_KEY="your-api-key"
Basic Valyu search example
Here is the smallest useful Valyu search call:
from valyu import Valyu
import os
client = Valyu(api_key=os.environ["VALYU_API_KEY"])
response = client.search(
query="Latest AI inference datacenter projects",
search_type="all",
max_num_results=5,
is_tool_call=True
)
print(response)
That is enough to search and get back relevant results.
A search result can include fields like:
- title
- URL
- source
- content
- description
- relevance score
- publication date
- source type
For an agent, the most important fields are usually:
titleurlcontentsourcepublication_date
Those fields let your agent reason from actual sources instead of unsupported model memory.
Build a tiny search agent
Now let’s wrap this in a simple function.
import os
from valyu import Valyu
client = Valyu(api_key=os.environ["VALYU_API_KEY"])
def search_agent(question: str, max_results: int = 5):
response = client.search(
query=question,
search_type="all",
max_num_results=max_results,
is_tool_call=True
)
sources = []
for result in response.results:
sources.append({
"title": result.title,
"url": result.url,
"source": result.source,
"content": result.content,
"publication_date": getattr(result, "publication_date", None),
})
return sources
if __name__ == "__main__":
question = "What are the latest AI inference datacenter projects?"
results = search_agent(question)
for i, source in enumerate(results, start=1):
print(f"\n[{i}] {source['title']}")
print(source["url"])
print(source["content"][:700])
This is not a full agent yet. But it is the search layer every useful agent needs.
Turn search results into LLM context
Most agents need search results formatted as context for a model.
Here is a simple formatter:
def format_sources_for_llm(sources):
formatted = []
for i, source in enumerate(sources, start=1):
formatted.append(
f"""
Source [{i}]
Title: {source["title"]}
URL: {source["url"]}
published: {source["publication_date"]}
Content:
{source["content"][:2000]}
"""
)
return "\n\n".join(formatted)
Now combine it with the search agent:
question = "What are the latest AI inference datacenter projects?"
sources = search_agent(question)
context = format_sources_for_llm(sources)
prompt = f"""
Answer the user's question using only the sources below.
User question:
{question}
Sources:
{context}
Instructions:
- Give a concise answer.
- Cite sources inline using [1], [2], etc.
- If the sources do not contain enough information, say so.
"""
print(prompt)
You can now pass this prompt to your preferred LLM.
The important part is that the LLM is no longer answering from memory alone. It is answering from retrieved sources.
Agent architecture
The architecture is simple:
This is the basic pattern behind many useful AI systems:
- research assistants
- due diligence agents
- finance agents
- scientific literature agents
- documentation agents
- market research bots
- competitive intelligence tools
- claim checkers
- monitoring agents
The domain changes, but the pattern stays the same.
Search API vs RAG
A common question is:
Why not just use RAG?
You should use RAG when your agent needs to search a known internal corpus.
Examples:
- company docs
- support tickets
- product manuals
- private research notes
- internal PDFs
You should use live search when your agent needs information that is:
- recent
- external
- changing
- public
- domain-specific
- not already in your vector database
In practice, strong agents often use both.
Internal knowledge → RAG
External knowledge → Search API
Long-form investigation → DeepResearch
Today’s example covers the external knowledge piece.
When this pattern is useful
This simple search-agent pattern is useful when you want to build any of the following.
Documentation agent - Search live docs and answer based on current API behavior.
Market research agent - Search the web, news, filings, and company pages.
Scientific assistant - Search papers, preprints, PubMed, clinical trials, or patents.
Finance agent - Search SEC filings, earnings, market data, and company fundamentals.
Monitoring agent - Ask “what changed this week?” across a topic, company, market, or research area.
Coding assistant with current context - Give your coding agent access to up-to-date documentation and examples.
A slightly better version
Let’s add a small final response object.
import os
from valyu import Valyu
client = Valyu(api_key=os.environ["VALYU_API_KEY"])
def valyu_search_agent(question: str, max_results: int = 5):
response = client.search(
query=question,
search_type="all",
max_num_results=max_results,
is_tool_call=True
)
sources = []
for index, result in enumerate(response.results, start=1):
sources.append({
"id": index,
"title": result.title,
"url": result.url,
"source": result.source,
"content": result.content[:1500],
"publication_date": getattr(result, "publication_date", None),
})
llm_context = "\n\n".join(
f"""
[{source["id"]}] {source["title"]}
URL: {source["url"]}
Published: {source["publication_date"]}
Content:
{source["content"]}
"""
for source in sources
)
return {
"question": question,
"sources": sources,
"llm_context": llm_context,
}
if __name__ == "__main__":
result = valyu_search_agent(
"What are the latest AI inference datacenter projects?"
)
print("Question:")
print(result["question"])
print("\nSources:")
for source in result["sources"]:
print(f'[{source["id"]}] {source["title"]} - {source["url"]}')
print("\nLLM Context:")
print(result["llm_context"])
Now you have a reusable search component that can be plugged into almost any agent framework.
What to build next
Once you have this basic search layer, you can extend it in a few directions.
Add an LLM: Use the search results to generate a final cited answer.
Add source filters: Search specific sources, domains, or datasets.
Add memory: Store previous questions and retrieved sources.
Add scheduled runs: Ask the same question every day or every week and detect changes.
Add DeepResearch: For complex topics, move from quick search results to long-form cited reports.
Add MCP: Connect search directly to Claude, ChatGPT, Cursor, Codex, or another MCP client.
We will cover these throughout the series.
FAQ
How do I add live web search to an AI agent?
Add a search tool that accepts the user’s question, retrieves relevant web results, and passes the result content and URLs into the model prompt. The model should answer using only the retrieved sources and cite them inline.
When should an AI agent use search instead of RAG?
Use search when the agent needs fresh, external, public, or domain-specific information that is not already in your internal knowledge base. Use RAG when the agent needs to retrieve from private or pre-indexed internal documents.
Can this work with ChatGPT, Claude, Cursor, or Codex?
Yes. You can expose Valyu search as a tool, call it from your backend, or connect it through MCP where supported.
What can Valyu search besides the web?
Valyu can search the live web and specialized sources across areas like academic literature, finance, SEC filings, patents, biomedical research, clinical trials, chemistry, drug discovery, legal data, cybersecurity, and more.
What is the difference between Valyu Search, Answer, and DeepResearch?
Valyu Search returns relevant source results. Valyu Answer returns a grounded answer from search. Valyu DeepResearch runs a longer research workflow that plans, searches, reads, and writes cited reports.
Why this matters for Hacktoberfest
Hacktoberfest this year is about open-source AI, agents, tools, and building.
Search is one of the most useful tools you can add to an agent.
A model can write, reason, and call tools. But without access to fresh and trustworthy sources, it cannot reliably know what is happening outside its own context window.
That is why Day 1 starts here. Before building complex agents, give them the ability to search and retrieve useful information.
Try this with Valyu
Valyu gives AI agents one API for live web search, domain-specific sources, cited answers, and DeepResearch.
Useful links to check out for more about Valyu:
- Docs: https://docs.valyu.ai/
- Platform: https://platform.valyu.ai/
- Search API: https://docs.valyu.ai/api-reference/endpoint/search
- MCP server: https://docs.valyu.ai/integrations/mcp-server
- Data sources: https://docs.valyu.ai/guides/datasources















