TinySearch logo

TinySearch

Organization
TinySuiteHQ

Shrink the web for your agents! Token-efficient, fast and fully local web search and research.

PublisherTinySuiteHQ
RepositoryTinySearch
LanguagePython
Forks
14
Stars
229
Available tools
0
Transport typestdio
Categories
LicenseMIT
Links
  • Connect tools to AI workflows

    TinySearch exposes MCP capabilities that can be used by compatible AI clients and agents.

  • 0 available tools

    Browse the callable actions below, including names and descriptions when provided by the server.

  • Ready-to-copy setup

    Use the installation snippets to configure this server in your preferred MCP client.

  • Open source signals

    229 stars and 14 forks from the linked repository.

TinySearch

Website PyPI version PyPI Downloads License: MIT Release Last Commit Docker Pulls Discord MCP Server FastAPI

TinySearch is a self-hosted web-research tool for AI agents. It searches the web, reads the best pages, removes low-value content, and returns compact evidence with source URLs.

Your model receives the useful passages instead of paying to process entire webpages.

TinySearch is part of TinySuite, a suite of focused tools designed to make agentic operations cheaper by minimizing token usage through smart retrieval, selection, and context-management techniques.

Choose a tier

TierUse it whenEntry pointSearch backend
1. Python libraryYou are building with TinySuite or Pythonpip install tinysuite-searchDDGS
2. One-command MCPAn MCP client should launch TinySearch for youuvx --from "tinysuite-search[server]" tinysearchDDGS
3. Docker + SearXNGYou want the full self-hosted stack and HTTP MCPdocker compose ... up -dBundled SearXNG

Tiers 1 and 2 need no search service. Tier 3 adds a dedicated SearXNG service, persistent model storage, and a network MCP endpoint. See the installation guide for the Docker setup.

The expensive part of agent research is context

A search result is not yet useful evidence. Agents often have to open several pages, ingest navigation and boilerplate, and spend paid input tokens deciding which passages matter.

TinySearch moves that work in front of the model:

mermaid
flowchart LR
    A[Question] --> B[Search and crawl]
    B --> C[Local hybrid reranking]
    C --> D[Compact evidence<br/>with source URLs]
    D --> E[Your agent]

That lowers cost in three ways:

  • Smaller model context. Only the best-ranked evidence chunks are returned, within a controlled evidence budget.
  • No metered search API required by default. TinySearch can search through DDGS without a paid search provider.
  • Local retrieval by default. ONNX embeddings and hybrid reranking run on your machine instead of creating embedding API charges.

Search broadly. Read locally. Pay the model only for the evidence that matters.

This is retrieval, not summarization: TinySearch selects the passages worth keeping with local BM25 and embedding rerank, it doesn't run a model over the page to rewrite or condense it. Every returned chunk is the original page text, unedited, so what you cite is what the page actually said. That keeps the pipeline fast and free to run locally, at the cost of not compacting as aggressively as a dedicated reduction model could. A learned reduction step is a direction we may explore later; it isn't part of TinySearch today.

Actual savings depend on the pages, evidence limits, client model, and provider pricing. TinySearch reduces the web content sent to the model; it does not control what the client does with that evidence afterward.

The cost panel uses an illustrative $3.00 per million input-token rate and excludes search, crawling, model output, and downstream agent use.

The naive baseline isn't a strawman product, it's the same pages TinySearch crawled for each query, fed to the model unfiltered, the way a generic "search, then fetch the page" tool (a plain web-search-plus-fetch loop, the kind built into most coding agents) would. Measured against the current recommended flow (search then scrape_urls) and counted on the actual MCP tool-result text, TinySearch's primary interface. Reproduce or rerun it yourself:

bash
python scripts/benchmark_token_savings.py --json-out report.json

Quick start

With uv installed, add TinySearch to any MCP client:

json
{
  "mcpServers": {
    "tinysearch": {
      "command": "uvx",
      "args": [
        "--python",
        "3.12",
        "--from",
        "tinysuite-search[server]",
        "tinysearch"
      ]
    }
  }
}

The client launches TinySearch over stdio when it needs it. No repository clone, hosted account, or paid search key is required.

Fast search starts without Chromium or an embedding model. The first scrape initializes Chromium; focused scraping also initializes the configured embedding model. Pre-warm both ahead of time if you will use those workflows:

bash
uvx --from "tinysuite-search[server]" tinysearch setup

The MCP and FastAPI servers keep the scraper browser warm between nearby requests, then close it after browser_idle_shutdown_seconds. Direct Python calls retain their short-lived, caller-owned lifecycle.

Prefer Docker, a remote MCP endpoint, or a source checkout? Follow the installation guide.

The MCP tools

ToolUse it when
search(items)You need fast, backend-ordered discovery without crawling or reranking; batch independent subquestions when useful
scrape_urls(items)You know one to five pages; each item may use * for its configured clean page-order token budget
browser_navigate / browser_actA page needs interaction before it can be read; see Browser automation
get_current_datetime()A question depends on the current date or time

TinySearch deliberately stays focused. It is a retrieval layer, not another agent, chat interface, hosted search product, or permanent web index.

See the complete MCP tool reference for parameters and response contracts.

What your agent gets

TinySearch does not spend another model call writing the final answer. The recommended flow is search for lightweight discovery, then scrape_urls for the pages worth reading.

Search returns structured JSON. Use one item for a simple lookup; add multiple items only for independent subquestions or source strategies. domains is a hard positive source restriction and accepts a domain plus its subdomains:

json
{"items":[{"query":"Form 8-K Tesla","domains":["sec.gov"]}]}

Each search item reports its own results and compact backend attempts. A zero result response is distinct from a blocked, unavailable, or invalid backend. scrape_urls returns selected Markdown evidence and separate related-link navigation candidates, each with independent configured token ceilings.

MCP still uses its standard JSON-RPC transport envelope, including protocol-level errors and optional structuredContent. Python and FastAPI keep their structured JSON contracts for applications that need to store, inspect, or transform the evidence.

How it works

  1. search returns backend-ordered titles, URLs, previews, upstream dates, and backend outcomes without starting Chromium or an embedding model.
  2. scrape_urls reads one to five known pages concurrently. Omit an item's query or use "*" to keep clean Markdown in page order within the configured token budget.
  3. Supply a focused item query when TinySearch should chunk and hybrid-rank that page before returning evidence.
  4. The browser_* tools step in only when scrape_urls can't reach the content because it needs interaction. See Browser automation.

Python library

TinySearch also works as a regular Python package:

bash
pip install tinysuite-search

Optional OpenTelemetry export

TinySearch emits vendor-neutral traces and metrics only when OpenTelemetry is explicitly configured. Normal library, MCP, and FastAPI behavior is unchanged when telemetry is not installed or not configured.

Install the optional exporter support for a standalone MCP server:

bash
uvx --from "tinysuite-search[server,telemetry]" tinysearch

The official Docker image includes the same optional support. Set standard OTel variables on either deployment; the common endpoint enables traces and metrics:

bash
OTEL_SERVICE_NAME=tinysearch \
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4318 \
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf \
tinysearch serve

http/protobuf is the default and grpc is also supported. Use OTEL_TRACES_EXPORTER=none, OTEL_METRICS_EXPORTER=none, or OTEL_SDK_DISABLED=true to disable telemetry. Standard resource, header, timeout, sampler, and signal-specific endpoint settings are passed through to the OpenTelemetry SDK; treat OTEL_EXPORTER_OTLP_HEADERS as a secret.

TinySearch exports operation/stage timing, outcomes, counts, backend state, browser use, token counts, and embedding model metadata. It never exports queries, URLs, domains, prompts, documents, snippets, request headers, credentials, configuration paths, raw errors, exception stacks, MCP arguments, or MCP results. Direct Python-library users configure their own OTel provider; TinySearch only auto-configures its standalone MCP and FastAPI server entry points.

python
import asyncio
from tinysearch import scrape_urls, search


async def main():
    results = await search([{"query": "Python async tasks"}])
    print(results["items"][0]["results"])

    page_url = results["items"][0]["results"][0]["url"]
    evidence = await scrape_urls([{
        "url": page_url,
        "query": "How does asyncio cancellation work?",
    }])
    print(evidence["results"])


asyncio.run(main())

The Python API returns stable, JSON-serializable results. search accepts one to five items and uses the configured per-item result limit. scrape_urls accepts a per-call max_tokens budget (4,000 by default); omit an item's scrape query or use "*" for page-order mode. Rendering structured evidence into an LLM prompt is explicit, so applications can store, inspect, transform, or budget the result first.

The optional FastAPI app mirrors these surfaces. POST /search accepts the same batch JSON contract. POST /scrape accepts one to five { "url", "query" } items and always returns structured per-item outcomes. POST /browser/navigate and POST /browser/act mirror the two MCP browser tools and return their accessibility view as { "result": "..." }. The app also exposes /health, /current_datetime, and read-only /config; configuration writes require explicit environment opt-in.

Search backends

TinySearch selects a web-search backend from config, so you can start with no search service and add one later without changing code.

  • "ddgs" (native default): queries the ddgs package's automatic backend selection in-process. No SearXNG deployment required.
  • "searxng" (Docker default): queries a self-hosted SearXNG instance. Falls back to ddgs on backend failure unless search_backend_fallback is set to false.
  • "duckduckgo": skips SearXNG and queries ddgs in DuckDuckGo-only mode.
  • "auto": tries SearXNG, then falls back to ddgs on any backend failure.

Set the BRAVE_SEARCH_API_KEY environment variable to add Brave's official Web Search API as a keyed fallback for the ddgs and duckduckgo backends. Brave is only consulted when the primary call errors or returns no results.

Full key reference, SearXNG JSON-output setup, and Compose details live in the configuration reference.

Browser automation

scrape_urls is a static fetch. It cannot see content that JavaScript renders after load, content behind a cookie interstitial, a "load more" control, or a client-side search UI. For those pages TinySearch drives a real browser.

This costs nothing extra to install. TinySearch already depends on Playwright through Crawl4AI and already installs its Chromium for scraping, so the browser tools reuse the same driver and the same browser: no second runtime, no second browser, no child process.

Two tools: browser_navigate and browser_act, the second folding look, click, type, wait_for, tabs, and close behind one action parameter.

That split is deliberate. MCP has no way to group or nest tools -- tools/list is flat and every schema is re-sent to the model on every request -- so seven separate browser tools would dominate the server's schema. Publishing the entry point as its own tool and folding one page session's lifecycle behind a dispatcher keeps the whole server's schema small across five tools.

There is no separate find tool, because finding is not a sibling of clicking -- it is a filter on the result. Both tools take one find argument, tried as a regex first (so "a|b" works directly) and falling back to a literal, case-insensitive substring match for text that isn't valid regex -- narrowing the return value to the matching nodes and their context instead of the whole tree. That is what lets one call both act and report: a click that reveals a table comes back as the table, so the agent never spends a second call narrowing the first one's answer. An earlier version split this into find and find_regex; that cost a wasted round trip whenever a model reached for alternation syntax on the plain-substring parameter and got "no matches" instead of a hint, so the two were merged.

The model reads a compact accessibility tree where each node carries a stable ref, names one, and TinySearch acts on it with genuine browser input events:

browser_navigate  -> url: "...", find: "Accept"   -> "- button \"Accept all\" [ref=e79]"
browser_act       -> action: "click", target: "e79", find: "Results"

Nothing synthesizes DOM events or invents CSS selectors, and click/type accept only a ref the model actually observed, never a raw selector.

Three deliberate choices:

  • No tool can execute code. There is no evaluate tool, and none that fills forms, uploads, or drags. A page that injects instructions into its own rendered text has nothing dangerous to reach for, because the capability is absent rather than discouraged.
  • find is the token lever, depth the fallback. Any call that returns a view takes find, cutting it to the matching nodes and their context. When no filter can name the target, depth returns a shallower but still valid tree rather than a truncated string -- on a large page ~700 characters versus ~33,000.
  • Cookies persist, sessions don't. Set browser_storage_state_path and a consent banner accepted once is not paid for on every later navigation. Sessions stay isolated, so no browser profile lock is taken and concurrent clients do not conflict. That file is a server-side path, never exposed to a model or over HTTP.

To turn the tools off entirely, set "browser_backend": "off"; they are then removed from the tool list rather than merely refusing to run.

External browser over CDP

To drive a browser you operate separately, set its Chrome DevTools Protocol endpoint. It is used by both the scrape pipeline and the browser tools:

json
{
  "browser_cdp_url": "http://browser:9222"
}

Server processes also accept TINYSEARCH_BROWSER_CDP_URL. The external browser owns its executable, profile, proxy, and fingerprint configuration; TinySearch does not select or install a particular browser backend.

Treat a CDP endpoint as privileged remote control of the browser. Keep it on a private network or loopback interface, require authentication when it crosses a host boundary, and do not expose port 9222 directly to the public internet. When TinySearch itself runs in Docker, localhost refers to the TinySearch container, so use an endpoint reachable from that container.

browser_backend, browser_cdp_url, and browser_storage_state_path are operator-managed and cannot be changed through the HTTP PUT /config endpoint, even when configuration writes are enabled. Set them in the startup environment (TINYSEARCH_BROWSER_BACKEND and friends) or the file selected by TINYSEARCH_CONFIG_PATH, then restart TinySearch. HTTP clients can continue updating other settings by omitting these fields from their partial update.

Proxies

To route crawling through a proxy, set browser_proxy_server:

json
{
  "browser_proxy_server": "http://gateway.proxyprovider.com:8000",
  "browser_proxy_username": "your-username",
  "browser_proxy_password": "your-password"
}

http, https, and socks5 proxy URLs are supported, matching what Playwright accepts. If your provider hands out several fixed gateway endpoints, comma-separate them and TinySearch round-robins one per new browser context:

json
{
  "browser_proxy_server": "http://gw1.proxyprovider.com:8000, http://gw2.proxyprovider.com:8000"
}

Use browser_proxy_bypass for a comma-separated list of hosts that should skip the proxy. These four fields are operator-managed like the CDP settings above -- settable via TINYSEARCH_BROWSER_PROXY_SERVER, TINYSEARCH_BROWSER_PROXY_USERNAME, TINYSEARCH_BROWSER_PROXY_PASSWORD, and TINYSEARCH_BROWSER_PROXY_BYPASS, or the config file, but never over HTTP.

Why TinySearch

  • No vendor in the loop. No TinySearch account, no required API key, no per-request billing, no analytics service or hosted scraped-data cache. The infrastructure you'd otherwise pay a search API for runs on your machine.
  • Source-grounded by construction. Every evidence chunk is the original page text, still attached to its originating URL, so a claim in your agent's answer traces back to one specific passage instead of stopping at "the vendor's model said this."
  • Built around token efficiency. Page selection and passage selection happen locally, before content enters model context.
  • Useful without paid infrastructure. DDGS search and local ONNX embeddings are the defaults.
  • Bring your own stack when needed. SearXNG and OpenAI-compatible embedding providers remain optional.
  • Works where agents already work. Use MCP over stdio, Streamable HTTP, Python, FastAPI, or Docker.

Part of TinySuite

TinySuite is a product suite built around one idea: agents should spend tokens on useful work, not operational overhead.

Each tool focuses on a different part of the agent workflow and uses targeted techniques to reduce unnecessary context before it reaches the model. TinySearch handles the web-research layer by turning pages into a small, ranked, source-grounded evidence packet.

Documentation

The README is the product overview. Detailed setup and operational material lives in the TinySuite documentation:

The repository also contains an annotated example configuration at configs/tinysearch_config.json.

When not to use TinySearch

TinySearch is intentionally lightweight. Use a commercial search API, persistent crawler, or full search index when you need:

  • guaranteed search coverage or an SLA
  • large-scale or scheduled indexing
  • long-term page storage and change history
  • enterprise observability and access controls

Development

bash
git clone https://github.com/TinySuiteHQ/TinySearch
cd TinySearch
python -m venv .venv
source .venv/bin/activate
pip install -e ".[server]"
python -m unittest discover tests

TinySearch supports Python 3.12 and newer. CI tests Python 3.12, 3.13, and 3.14 across Linux, macOS, and Windows.

Entrypoints

  • tinysearch.search and tinysearch.scrape_urls: structured Python API
  • tinysearch.get_current_datetime: structured UTC date and time
  • tinysearch.to_prompt: pure structured-evidence prompt renderer
  • tinysearch mcp: stdio MCP server (also the no-argument default)
  • tinysearch serve: Streamable HTTP MCP server
  • tinysearch.servers.fastapi_server:app: optional FastAPI application

Community

Questions, ideas, and bug reports are welcome:

Privacy and license

TinySearch reads public pages and returns selected excerpts to the calling client. Search, crawling, local embeddings, and reranking can run without sending page content to an embedding provider. If you choose an OpenAI-compatible embedding backend, that provider receives the text sent for vectorization.

TinySearch is available under the MIT License. Downloaded model weights remain subject to their respective model-card licenses. See NOTICE for third-party distribution details.

Use TinySearch MCP with multiple AI models

TypingMind connects MCP tools at the workspace level, so once TinySearch is connected, you can use it with different AI models in TypingMind instead of setting it up separately for each model. This MCP runs locally through the TypingMind MCP connector on your device.

Setup guide to use the local connector

Use this when the MCP server needs access to local files, apps, or private resources on your computer.

1

Open the MCP settings

In TypingMind, go to Settings, Advanced Settings, then Model Context Protocol and choose Setup Connector.

  1. Open TypingMind in your browser.
  2. Click the Settings icon.
  3. Go to Advanced Settings.
  4. Open the Model Context Protocol section.
  5. Click Setup Connector and choose This Device.
TypingMind MCP connector setup screen with This Device selected
2

Run the connector command

Choose This Device, copy the command from TypingMind, and run it in Terminal. Keep the process running while you use MCP.

  1. Copy the setup command shown by TypingMind.
  2. Open Terminal on macOS or Windows Terminal on Windows.
  3. Paste and run the command.
  4. Approve the package install if Terminal asks you to proceed.
  5. Keep the Terminal window running while using MCP tools.
3

Add TinySearch as a server

When the connector status is Ready, click Edit Servers and paste the MCP server configuration.

  1. Wait until the connector status shows Ready.
  2. Click Edit Servers.
  3. Paste the TinySearch MCP server configuration.
  4. Save the server list.
  5. Refresh if you want to confirm the connector is still ready.
TypingMind MCP settings showing active server and Edit Servers button
{
  "mcpServers": {
    "tinysearch": {
      "command": "npx",
      "args": [
        "-y",
        "<mcp-server-package>"
      ]
    }
  }
}
4

Use it across models

Save the server list, open Plugins, enable the TinySearch MCP tools, then select any supported AI model in TypingMind and use the tools in chat or assign them to an AI agent.

  1. Open the Plugins page in TypingMind.
  2. Enable the TinySearch MCP tools.
  3. Start a chat and choose the AI model you want to use.
  4. Use the MCP tools in chat or assign them to an AI agent.
  5. Switch to another AI model whenever needed without reconnecting MCP.
TypingMind chat using enabled MCP tools with a selected AI model
Can you use TinySearch to help me with this task?
TinySearch
Sure. I read it.
Here is what I found using TinySearch.

Frequently asked questions

What is the TinySearch MCP server used for?

TinySearch is an MCP server that lets compatible AI clients connect to external tools and context. In TypingMind, you can add this MCP server once and make its tools available in your AI workspace.

Can I use TinySearch MCP with multiple AI models in TypingMind?

Yes. TypingMind connects MCP tools at the workspace level, so you can use TinySearch with different AI models such as Claude, ChatGPT, Gemini, or other models you have configured in TypingMind without setting up the MCP server separately for each model.

Why use TinySearch MCP with TypingMind?

TypingMind is one of the best frontends for LLM chat because it brings multiple AI models, prompts, plugins, AI agents, API keys, and MCP tools into one workspace. With TinySearch connected, you can use its MCP tools across your preferred models while keeping your chat workflow organized in TypingMind.

How do I connect TinySearch MCP to TypingMind?

TinySearch runs through the TypingMind local MCP connector. This is best when the MCP server needs access to local files, desktop apps, command-line tools, or private resources on your computer.

What tools does TinySearch MCP provide in TypingMind?

TinySearch exposes MCP capabilities that can be enabled from the TypingMind Plugins page and used in chat or assigned to AI agents.

Do I need to share my API keys with TypingMind to use TinySearch MCP?

No. TypingMind is local-first and lets you keep your model providers, API keys, prompts, and MCP configuration under your control. If TinySearch requires authentication, add the required headers, OAuth settings, or local configuration for that MCP server when you create the connection.

Related MCP Servers

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇