Research · BSc thesis · SUPSI

Local RAG for medical guidelines

Optimization of a RAG System with a Local LLM for Evidence-Based Medical Guideline Retrieval. Clinical questions can contain patient data, so nothing leaves the building: the model, the embeddings and the vector store all run on local machines.

From question to cited answer

Follow the evidence through the local pipeline. Select a step to take a closer look.

06/06

The answer points back to its source.

Illustrative walkthrough. Passage scores and citation references are simulated; no live model is running.

94/99

correct answers, internal LLM judge

90/99

correct, stricter external judge

98/99

expected guideline retrieved

66 → 81

local score gained by reranking

The pipeline

  1. 1Ingest

    PDF parsing and chunking with document and page metadata, plus OCR of figures.

  2. 2Route

    Dense retrieval with routing per guideline, and a hard filter when the question names a document.

  3. 3Rerank

    A cross-encoder reorders a deep pool of 20 passages. This was the single biggest gain.

  4. 4Answer

    Targeted prompting; every answer cites the guideline and the page it comes from.

Corpus & hardware

44 clinical guidelines in nephrology and internal medicine, 3,236 PDF pages. Full run of 99 questions in 3.01 hours, about 109 s per question including grading.

  • NVIDIA Jetson AGX Orin
  • qwen3:32b via Ollama
  • bge-m3 embeddings
  • PostgreSQL + pgvector
  • Cross-encoder reranker
  • Web front end

How it was evaluated

LLM-as-judge across six configurations, inter-judge agreement checks, grounding analysis (verbatim 6-gram coverage), expert clinical review, ablations and cost/latency measurements.

What it did not reach. The 95% accuracy target was not reached. Before any clinical use the thesis calls for better grounding in the cited evidence, an independent clinical benchmark and a new expert review.

Supervisor Vanni Galli · Host company Alwicom SA · 25 August 2026

Coursework & other ML

  • Citation link prediction on DBLP

    LightGBM, XGBoost, Random Forest, MLP and a Graph Convolutional Network, with feature engineering, a data-leakage audit, SHAP, uncertainty estimates and stratified edge subsampling.

    • Graph ML
    • GCN
    • SHAP
  • CYP2C19 inhibition prediction

    Cheminformatics with RDKit: ECFP4 fingerprints, physico-chemical descriptors, SHAP, applicability domain and Williams plot.

    • RDKit
    • Cheminformatics
  • Lipophilicity prediction

    RDKit descriptors, Spearman correlation matrices, permutation feature importance and forward feature selection.

    • Regression
    • Feature selection
  • Autonomous robot navigation

    A DJI RoboMaster EP navigating on its own in a CoppeliaSim simulation.

    • Robotics
    • Simulation
  • Penetration testing lab

    Metasploit, antivirus evasion and obfuscation techniques, in a controlled teaching environment.

    • Security

Jump anywhere

Search pages, profiles and actions