Selected work

Azure RAG engineering · 2026

University finance AI assistant

An evidence-led assistant for exploring public university financial reports, built around measurable retrieval, page-level provenance and safe abstention.

  • Python
  • Azure AI Search
  • Microsoft Foundry
  • RAG evaluation

Challenge

Financial reports combine narrative, tables, changing periods and similar figures. The system must retrieve the right evidence, preserve scope and units, cite the original PDF and decline questions the indexed documents cannot answer.

My role

I designed the staged ingestion and evaluation pipeline, implemented Azure hybrid search and grounded generation, and separated the public interface from credentials and model access.

Outcome

The reviewed ten-question baseline achieved 100% Recall@5 and 0.825 MRR@5. A separate ten-question negative set produced 100% correct, citation-free abstention.

At a glance

1.00hybrid Recall@5
0.825hybrid MRR@5
10/10correct abstentions

Interactive evidence

Ask the annual report.

Answers are restricted to the selected public document. Supported answers include page-level citations; unsupported questions should be declined.

Open source PDF

Suggested questions

Assistant

Select a suggested question or ask about the financial report.

System view

The workflow.

A simplified view of the stages and boundaries that shape the project.

  1. 01Public PDF
  2. 02Validated chunks
  3. 03Azure embeddings
  4. 04Hybrid retrieval
  5. 05Grounded answer
  6. 06Page citations

Approach

Decisions that shaped the work.

01

Evaluate retrieval before generation

Reviewed question sets measure Recall@k and MRR independently, so a plausible language-model answer cannot hide missing evidence.

02

Make provenance part of the contract

Every answer must cite retrieved chunk IDs that map back to a public source document and exact PDF page.

03

Test when the system should refuse

A separate negative dataset checks that unsupported questions produce an abstention without decorative or misleading citations.

Findings

What the evidence says.

  • Hybrid BM25 and vector retrieval found a relevant chunk in the top five for all ten reviewed questions.
  • The first negative-dataset draft exposed a labelling error: the supposedly missing Moody’s rating was present on page 83.
  • The public serving boundary keeps Azure credentials out of GitHub Pages and supports an immediate cost-control switch.

Engineering reflection

The current evaluation is intentionally small and development-reviewed. Before production use, I would add independent domain review, multi-document regression tests, persistent distributed rate limiting and operational monitoring.