
project / sneakdex
SneakDex
Built a high-performance, distributed semantic search engine using Python, Go, Rust, and modern cloud-native tools. Delivers lightning-fast, relevant results at scale.
Building SneakDex: A Modern Distributed Search Engine for the Web-Scale Era
How I built an enterprise-grade search engine with microservices, semantic search, and sub-second response times
The Challenge
In an era where billions of web pages compete for attention, traditional search engines often struggle with three fundamental challenges: scale, semantic understanding, and real-time processing. We set out to solve these problems by building SneakDex—a high-performance, distributed search engine designed from the ground up for modern web-scale content discovery.
What started as an ambitious project to understand search infrastructure evolved into a production-ready system capable of processing thousands of pages per minute while delivering sub-second search results. Here's how we did it.
Architecture Overview: The Microservices Approach
The heart of SneakDex lies in its microservices architecture. Rather than building a monolithic system, we decomposed the search pipeline into four specialized services, each optimized for its specific task:
The Crawler: Discovering the Web
Built with Go and the powerful Colly framework, our crawler is the engine's eyes and ears. It efficiently discovers and fetches web content while respecting website boundaries and preventing abuse.
Why Go? We needed raw performance and excellent concurrency primitives. Go's goroutines allow us to crawl thousands of pages simultaneously without the overhead of traditional threading models.
Key innovations:
- Redis-based distributed queue: Multiple crawler instances coordinate through Redis, ensuring no URL is processed twice
- Intelligent rate limiting: Respects robots.txt and implements exponential backoff to be a good web citizen
- Real-time streaming: Discovered content flows immediately to Kafka for downstream processing
- Security-first design: Filters private IP addresses, validates domains, and enforces content size limits
The crawler processes 1000+ pages per minute per instance, with linear scalability—spin up more instances, get proportionally more throughput.
The Parser: Extracting Signal from Noise
When raw HTML arrives from the crawler, our Rust-powered parser transforms it into structured, searchable data.
Why Rust? HTML parsing is memory-intensive and error-prone. Rust's memory safety guarantees and zero-cost abstractions give us both speed and reliability—no garbage collection pauses, no memory leaks.
What it does:
- Extracts clean text from HTML chaos (goodbye script tags and inline styles)
- Identifies page language automatically using the whatlang library
- Pulls out metadata: titles, descriptions, headings (H1-H6)
- Catalogs links and images for comprehensive indexing
- Validates content quality and filters out noise
The parser outputs pristine JSON payloads with normalized text, ready for semantic analysis and indexing.
The Indexer: Making Content Searchable
This is where the magic happens. Our Python-based indexer leverages cutting-edge machine learning to create both semantic and traditional search indexes.
The dual-indexing strategy:
Vector indexing with Qdrant: We use Sentence Transformers to generate 384-dimensional vector embeddings that capture semantic meaning. "Best Italian restaurants" and "top pizza places" become mathematically similar, even though they share no common words.
Sparse indexing with PostgreSQL: Traditional full-text search provides precise keyword matching and handles queries where exact terms matter.
Why both? Hybrid search combines the semantic understanding of vectors with the precision of keyword matching, giving users the best of both worlds.
Processing pipeline:
- Consumes parsed content from Kafka in real-time
- Generates embeddings in configurable batches (default: 50 documents)
- Stores vectors in Qdrant for blazing-fast similarity search
- Indexes metadata and text in PostgreSQL with tsvector for full-text search
- Processes images with captions for text-to-image search capabilities
Performance: The indexer processes batches concurrently, handles malformed messages gracefully, and scales horizontally to meet enterprise workloads.
The App: Bringing It All Together
The user-facing application, built with Next.js 15, is where all the pieces converge into a seamless search experience.
The hybrid search algorithm:
Final Score = (0.75 × Vector Similarity) + (0.25 × Full-Text Score) + Domain Boost
This weighted combination prioritizes semantic relevance while ensuring exact matches aren't lost. Domain boosting helps surface authoritative sources based on query characteristics.
Two search modes:
Traditional web search: Queries both the vector database and PostgreSQL, then intelligently fuses results using our scoring algorithm. Multiple fallback mechanisms ensure high availability with a vector search to payload fallback to PostgreSQL chain.
Text-to-image search: Pure vector search matches text queries against image embeddings, enabling natural language image discovery: "sunset over mountains" finds relevant images without manual tagging.
Performance optimizations:
- Multi-layer caching: In-memory embeddings plus Redis persistence
- Smart TTL management: Frequently accessed results stay hot
- HuggingFace fallback: Local ML models with cloud fallback for reliability
- Cache hit rate: 80%+ reduces database load dramatically
The result? Sub-second response times even with millions of documents indexed.
The Technology Stack: Best Tool for Each Job
We embraced polyglot programming, choosing languages optimized for each service's requirements:
Crawler (Go): Concurrency primitives, network performance
Parser (Rust): Memory safety, zero-cost abstractions
Indexer (Python): ML ecosystem, Sentence Transformers
App (TypeScript/Next.js): Full-stack React, API routes, modern developer experience
Infrastructure:
- Kafka: Real-time event streaming between services
- Redis: Distributed queue and caching layer
- Qdrant: High-performance vector database
- PostgreSQL: Robust full-text search and metadata storage
- Docker and Kubernetes: Container-first deployment
Security: Building Trust from the Ground Up
Search engines touch the entire web, making security paramount. SneakDex implements multiple layers of protection:
- RFC 1918 private IP filtering: Never access internal networks
- Domain validation: Whitelist/blacklist support prevents abuse
- Content size limits: Protects against DoS via oversized payloads
- Request timeouts: Prevents hanging on slow/malicious servers
- Transparent User-Agent: Identifies ourselves to webmasters
- Secrets management: Environment-based, never committed to code
- Input sanitization: Query validation prevents injection attacks
- Encrypted connections: All database communications secured
Performance by the Numbers
Real-world metrics from our production deployment:
- Crawling: 1,000+ pages/minute per instance
- Indexing: 50 documents per batch (configurable)
- Search latency: Sub-second response times
- Vector search: Sub-millisecond similarity lookups
- Cache performance: 80%+ hit rate
- Uptime: 99.9% with comprehensive fallbacks
- Concurrent users: Horizontal scaling supports unlimited growth
Lessons Learned
Microservices Pay Off at Scale
The initial complexity of managing multiple services proved worthwhile. Each service can scale independently based on bottlenecks, and failures are isolated.
Hybrid Search Beats Pure Vector Search
While semantic search is powerful, combining it with traditional full-text search captures both intent and precision.
Rust for Data Processing is Brilliant
The parser never crashes, never leaks memory, and processes millions of pages without supervision. Rust's guarantees are real.
Caching is Your Best Friend
Our 80%+ cache hit rate means the database only handles 20% of queries. This single optimization made sub-second search possible.
Observability from Day One
Prometheus metrics and structured logging saved countless hours during debugging and optimization. Build it in from the start.
What's Next?
SneakDex continues to evolve. On our roadmap:
- Personalization: User-specific ranking signals
- Multi-language support: Expanded language detection and localization
- Advanced image search: OCR and object detection
- Real-time crawling: Event-driven updates for breaking news
- GraphQL API: More flexible querying for developers
Open Source and Contributing
SneakDex is MIT licensed and available on GitHub. We welcome contributions, whether it's bug reports and fixes, feature suggestions, documentation improvements, or performance optimizations.
Star us on GitHub: https://github.com/Sneakyhydra/SneakDex
Conclusion: Search for the Modern Web
Building SneakDex taught us that modern search engines need more than just indexing—they need semantic understanding, real-time processing, and architecture that scales. By combining microservices, machine learning, and battle-tested infrastructure, we created a search engine ready for web-scale challenges.
The journey from concept to production was filled with technical challenges, architectural decisions, and performance optimizations. But the result is a system that proves sophisticated search infrastructure is within reach of any motivated team.
The web deserves better search. We're building it.
Have questions about our architecture or want to discuss search infrastructure? Drop a comment below or reach out on GitHub. We love talking search!
Built with love for the open web