Back to work
NLP / RAGCase study

Code & Clause

A retrieval-augmented assistant that answers Nigerian tech-law and NITDA policy questions.

Stack

FastAPIStreamlitRAGsentence-transformersGoogle GeminiPython

Highlights

  • PDF ingestion + voice input
  • Grounded retrieval with citations
View source

Regulatory text is long, dense and rarely searchable in a useful way. Code & Clause ingests Nigerian technology policy and NITDA documents, embeds them, and answers natural-language questions with grounded citations instead of guesswork. It accepts PDF uploads and voice input, so a founder can interrogate a policy document conversationally.

The problem

Nigerian technology regulation is spread across long PDF policy documents that founders and engineers rarely read end to end. Keyword search over them is close to useless, and a general-purpose chatbot will confidently invent provisions that do not exist.

Approach

  1. Built an ingestion pipeline that chunks uploaded policy PDFs and embeds them with sentence-transformers, so retrieval works over semantic meaning rather than exact wording.

  2. Grounded every answer in retrieved passages via a RAG pipeline over Google Gemini, keeping responses tied to source text rather than model recall.

  3. Exposed the system through a FastAPI service with a Streamlit interface, adding voice input so questions can be asked conversationally.

Outcome

A working assistant that turns static regulatory PDFs into something you can question directly, with answers traceable back to the passage they came from.