CareerCompass AI
Architecting an End-to-End AI Recruitment Platform with Change Data Capture (CDC), Kafka, Vector Search, and RAG
Role
Tech Lead / Architect
Evaluation
9.8 / 10 (Rank #1)
Precision
95% Retrieval Accuracy
Latency
< 150ms Search
1. Problem Statement
Traditional recruitment engines rely heavily on static SQL queries and keyword-based filters. When a job description asks for a "Distributed Systems Engineer familiar with Kafka," a qualified candidate who describes their experience as "architecting event-driven pipelines using publish-subscribe queues" is often completely omitted from initial search passes.
Furthermore, recruiters spend countless hours manually cross-referencing resumes against dense requirements. From a systems standpoint, performing real-time vector embedding generation and similarity calculations inside synchronous HTTP request-response cycles causes severe latency spikes and risks cascading timeouts.
2. System Architecture & Data Pipeline
To decouple transactional writes from heavy indexing workloads, I designed an event-driven Change Data Capture (CDC) architecture:
[ Client Web App ] ─── HTTP REST / JSON ───► [ FastAPI Core API Services ]
│
▼ (ACID Transactional Writes)
[ PostgreSQL Database ]
│
▼ (Write-Ahead Log / WAL)
[ Debezium CDC Connector ]
│
▼ (Event Stream: inserts / updates)
[ Apache Kafka Clusters ]
│
▼ (Asynchronous Consumer Groups)
[ Background Workers ]
├── OpenAI Embedding Pipeline
└── Batch Vector Normalization
│
┌───────────────────┼───────────────────┐
▼ ▼ ▼
[ Qdrant Vector DB ] [ Weaviate Engine ] [ Elasticsearch ]
(Dense Embeddings) (Hybrid Schemas) (Keyword / BM25)
│ │ │
└───────────────────┴───────────────────┘
│
▼
[ Recruiter Matching & RAG Engine ]
│
▼
[ OpenAI Assistants API ]
│
▼ (Charts & Metrics Artifacts)
[ MinIO Object Storage ]CareerCompass AI — System Architecture Diagram
Microservices, CDC Streaming, Kafka Event Bus, and Vector RAG Pipeline
Write Path Decoupling: When an applicant updates their profile or a recruiter posts a new role, the write commits instantaneously to PostgreSQL without stalling for vector embeddings.
Debezium CDC: Debezium tails PostgreSQL's Write-Ahead Log (WAL), streaming atomic event envelopes into topic partitions in Kafka. This guarantees that no indexing events are lost, even in the event of worker crashes.
Asynchronous Indexing: Dedicated worker pools consume Kafka events, invoke the OpenAI Embedding model with exponential backoff and rate-limiting, and upsert vectors into Qdrant and Weaviate.
3. Key Engineering Highlights
1. RAG & LLM Query Expansion
Queries are expanded via lightweight LLM prompts to extract domain synonyms, required competencies, and seniority levels. The resulting dense query vectors achieved 95% retrieval precision over raw lexical search.
2. Real-Time CDC Streaming
Eliminated manual dual-write patterns and two-phase commits. By reading directly from PostgreSQL WAL logs, the system guarantees zero dual-write inconsistencies between the relational store and vector databases.
3. Recruiter Semantic Matcher
Built a multi-criteria vector scoring engine that computes cosine similarity between candidate skill vectors and job requirements, returning rank-ordered applicant shortlists in sub-150ms.
4. AI Analytics & MinIO Storage
Integrated the OpenAI Assistants API with custom code-interpreter capabilities to generate dynamic visualizations and conversion metrics, persisting generated chart artifacts safely to a private MinIO S3 cluster.
4. Engineering Challenges & Solutions
Challenge 1: Event Ordering & Race Conditions During Profile Updates
Rapid successive edits to candidate profiles caused race conditions where older updates could overwrite newer vectors if processed out of order across worker threads.
Solution: Partitioned Kafka topics by user_id, ensuring that all state mutations for a specific user were routed to the same partition and consumed strictly in FIFO order.
Challenge 2: OpenAI API Rate Limits & Cost Control
High-volume resume ingestion quickly triggered HTTP 429 rate limits and threatened to incur excessive API costs.
Solution: Built an in-memory Redis embedding cache keyed by SHA-256 text hashes, deduplicating identical job descriptions and skill summaries. Implemented a token-bucket rate limiter within the background worker service.
Challenge 3: Hybrid Search Balance (Keywords vs Semantics)
Pure vector search occasionally omitted strict requirements (e.g., exact visa status or mandatory security clearance).
Solution: Adopted a two-stage hybrid retrieval strategy: first filtering candidates using Elasticsearch BM25 / Boolean facets, then re-ranking the top candidate pool with Qdrant vector similarity scores.
5. Results & Key Learnings
The capstone project was presented to faculty and industry reviewers, receiving a score of 9.8 / 10 and ranking #1 across all graduate capstones.
Leadership Takeaways: Serving as the technical lead for 6 engineers reinforced the value of strict interface contracts and early schema governance. Setting up Kafka topic definitions and Protobuf / Pydantic schemas upfront allowed the frontend, backend, and data pipelines to be developed concurrently without blocking dependencies.
Project Repositories
Explore the multi-repository architecture on GitHub:
