Back to work
NLP / RAGCase study

Chat With Docs

A credit-assessment chatbot grounded in API documentation, with real auth and a vector store.

Stack

FastAPISQLAlchemyAlembicPostgreSQLPineconeGroq Llama-3.3-70BJWT Auth

Highlights

  • Managed migrations via Alembic
  • JWT-authenticated endpoints
View source

A production-shaped RAG service rather than a notebook demo: JWT-authenticated, backed by PostgreSQL with Alembic migrations for relational state, and Pinecone for vector retrieval. It answers credit-assessment questions grounded in API documentation, served by Groq-hosted Llama 3.3 70B for low-latency inference.

The problem

Most retrieval demos stop at a notebook: no authentication, no schema management, no separation between relational and vector state. That is not deployable against real credit data.

Approach

  1. Split persistence deliberately — PostgreSQL with SQLAlchemy and Alembic for relational records and migration history, Pinecone for embedding retrieval.

  2. Put JWT authentication in front of the endpoints so document access is scoped rather than open.

  3. Served inference through Groq-hosted Llama 3.3 70B, trading self-hosting for materially lower response latency.

Outcome

A RAG backend with the operational parts intact — migrations, auth and separated stores — that answers credit-assessment questions grounded in the documentation it was given.