Clean search MCP across Yandex, Bing and DuckDuckGo
clean-search-mcp
A spam-filtered search MCP server: Yandex, Bing and DuckDuckGo, a 176K-domain blocklist and result quality scoring
Low risk
We rate an entry low when it mostly gives the agent instructions and reference material.
Why this level
- The server only reads public search results and page content
- Scraping HTML results instead of an official API risks a User-Agent block on Yandex's or Bing's side
Install
Manual install
pip install -r requirements.txt && python main.pyQuick start from the README; dependencies are lightweight, Playwright is optional for deep mode.
This is third-party code. Review the repository files before installing.
What it does
The server provides a clean_search tool that queries three engines with automatic fallback: it works through the sources listed in config.SEARCH_PROVIDERS, scraping Yandex HTML results at yandex.ru with no official API, scraping Bing HTML too, and using the duckduckgo_search library for DuckDuckGo. Results pass three filtering layers: a 176K-domain blocklist aggregated from 25-plus community sources, content rules, and a 0-to-1 quality score where official docs outrank tutorials and garbage content scores zero. It extracts full page text via trafilatura and selectolax, caches for 6 and 24 hours, and has a deep mode with a Playwright fallback for JS-heavy pages. Users can add domains to a personal blocklist and report bad results.
Who it is for. For agent developers who want spam- and content-farm-free search with Russian-language coverage via Yandex.
Good fit when
- You want one search across Yandex, Bing and DuckDuckGo with fallback between them
- Filtering spam and content farms before results reach the LLM matters
- You want extracted page text right away, not just a link and snippet
Not a fit when
- You need to stay strictly within Yandex's official API: Yandex results here come from HTML scraping, not the Search API
- You need production-grade stability: scraping can break if the markup changes or the User-Agent gets blocked
Example request
Find five high-quality sources about configuring Nginx, no content farms, and give me their full textLimitations
Yandex and Bing search is implemented by scraping the HTML results page with no official key, rather than via the Search API, so markup changes or blocks can break the parser at any time. Some source code and comments are in Chinese. Deep mode with Playwright requires installing a separate heavy package. clean_search caps out at 10 results per call.
How to disable. Remove the clean-search entry from mcpServers and stop the python main.py process.
MCP
- Transport
- stdio
- Authentication
- not required
Security check
- The server only reads public search results and page content
- Scraping HTML results instead of an official API risks a User-Agent block on Yandex's or Bing's side
README in short
The README lists the features: three engines with fallback, a 176K-domain blocklist from 25-plus sources, three-layer filtering, content extraction, quality scoring, LRU caching, a user blocklist and optional Playwright for heavy pages. It shows quick-start commands, an MCP client config, local testing, and the clean_search function parameters.
FAQ
Does it use the official Yandex API?
No, Yandex is queried by scraping the results page HTML; no official key is needed, but stability is not guaranteed.
Can I add my own domain to the blocklist?
Yes, add_user_blacklist adds a domain to a personal list, and report_bad_result blocks a URL after a complaint.
Related
CLI, skill and Python library for browser control: the agent clicks, fills forms and reads pages over CDP
A skill that gathers the last 30 days of discussion on a topic from Reddit, X, YouTube, HN, Polymarket and GitHub into one brief
The Chrome DevTools team's official MCP server: the agent drives a live Chrome, reads network and console, and records performance traces
A fast Rust CLI for agent browser automation: accessibility snapshots with element refs, an MCP server and skills