Haseeb Sagheer
Live

Ask Haseeb AI

Ask Haseeb AI is a retrieval-augmented generation (RAG) assistant built by Haseeb Sagheer. It ingests documents from Google Drive, stores embeddings in Pinecone, and answers questions through a FastAPI backend and a chat interface. It is live at ask.haseebsagheer.com and the code is public.

Ask Haseeb AI chat page answering the question What has Haseeb built, with options for visitor type and answer length
Screenshot of ask.haseebsagheer.com.

Overview

Live

Live means you can use it now at ask.haseebsagheer.com. It was first built in 2025 and rebuilt in October 2026 with a new interface, streaming answers and a new knowledge base. The source code is public on GitHub, and there is a recorded walkthrough of the first version further down this page.

Try it at ask.haseebsagheer.com

Source code on GitHub

Ask Haseeb AI is an assistant that answers questions about my skills, projects and experience. It does not rely on what a language model happens to know. It retrieves the relevant parts of my own documents and writes the answer from those.

Documents live in a Google Drive folder. When a new file appears, the system picks it up, cleans it, splits it into overlapping chunks of about 1,000 characters, embeds each chunk, and stores the vectors in Pinecone. A question is embedded the same way, the closest chunks are retrieved, and an OpenAI chat model writes the answer from that context.

Around that sits a FastAPI backend with REST endpoints, a simple chat front end, and a production deployment on a VPS behind Nginx with HTTPS. It is the project where I built every layer of a retrieval system myself.

The problem

Recruiters and clients ask the same questions about skills, projects and experience. The assistant answers them from the actual documents instead of a static page.

Who it is for

Recruiters, clients and collaborators who want answers about my work without reading every page.

What Ask Haseeb AI does

  • Detects new files in Google Drive and processes them automatically
  • Handles PDF, Markdown, HTML, plain text and Google Docs
  • Chunks text at about 1,000 characters with overlap to keep context
  • Streams answers and keeps the context of follow-up questions
  • Stores OpenAI embeddings in Pinecone
  • Retrieves relevant chunks and generates an answer with an OpenAI chat model
  • Serves queries through REST endpoints and a chat interface

How it works, step by step

Ask Haseeb AI flow5 steps. Select one to pause.

Engineering notes

Ingestion
The Google Drive API detects new files and triggers processing. PDF, Markdown, HTML, plain text and Google Docs are supported. Edited files are re-indexed and deleted files are removed.
Chunking
Documents are split into chunks of about 1,000 characters with overlap, so meaning carries across chunk boundaries.
Retrieval
OpenAI embeddings (3072 dimensions) are stored in Pinecone and searched for each question.
Serving
A FastAPI backend exposes REST endpoints for queries and background jobs. Answers are streamed as they are written. It runs on a VPS behind Nginx over HTTPS.
Index management
The index is recreated automatically and document metadata is managed alongside it.

Watch the walkthrough

A recorded walkthrough from architecture to deployment. It loads from YouTube when you press play.

Ask Haseeb AI project walkthrough video

My role

I built the whole system: ingestion, preprocessing, embeddings, the RAG chain, the API, the chat front end and the deployment.

What is next

  • Role-based access for public and private data
  • Query logs and document usage statistics
  • Translation for wider access
  • A React or Next.js front end
Status
Live
Type
Automation tool
Built with
Python, FastAPI, OpenAI, Pinecone, Google Drive API, Nginx
Built by
Haseeb Sagheer, solo
Page updated

Questions about Ask Haseeb AI

What is Ask Haseeb AI?

A RAG assistant that answers questions about Haseeb Sagheer’s skills, projects and experience from his own documents.

Is the demo live?

Yes, at ask.haseebsagheer.com. The code is on GitHub.

Which models and tools does it use?

OpenAI embeddings and an OpenAI chat model for answers, Pinecone for vector search, and FastAPI for the backend.

Where does its knowledge come from?

From documents in a Google Drive folder, which are ingested automatically when new files appear.

Need something like Ask Haseeb AI?