Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions clinical-note-similarity-py/.env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
DOCUMENTDB_URI=mongodb://<username>:<password>@localhost:10260/?tls=true&tlsAllowInvalidCertificates=true&authMechanism=SCRAM-SHA-256
DOCUMENTDB_DATABASE=clinicaldb
DOCUMENTDB_COLLECTION=notes
OLLAMA_BASE_URL=http://127.0.0.1:11434
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
FLASK_PORT=5001
NEAREST_NEIGHBORS=5
EMBEDDING_DIMENSIONS=768
9 changes: 9 additions & 0 deletions clinical-note-similarity-py/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
.env
__pycache__/
*.pyc
*.pyo
*.egg-info/
dist/
build/
.venv/
venv/
211 changes: 211 additions & 0 deletions clinical-note-similarity-py/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,211 @@
# Clinical Note Similarity Explorer — Python

> **Important:** All clinical notes in this sample are fictional and de-identified. This tool is for demonstration purposes only and is not intended for use with real patient data.

A Flask web application that stores de-identified clinical notes as documents with vector embeddings in **DocumentDB OSS**. Clinicians and researchers can search for similar cases using natural language descriptions — finding notes by clinical meaning rather than exact keyword matches. Built entirely on open-source tools with no cloud accounts required.

## Use cases

- Search by symptoms: *"patient with acute chest pain radiating to the left arm with ST elevation"*
- Search by presentation: *"shortness of breath not responding to bronchodilators with wheeze"*
- Search by findings: *"sudden onset facial droop, arm weakness, and slurred speech"*
- Research patterns: *"abdominal pain with right lower quadrant tenderness and fever"*

## How it works

```
Clinician Query (natural language)
Embed with nomic-embed-text (768-dim)
DocumentDB $search ──── cosine similarity ────► Similar Cases
(vector-ivf index) by clinical meaning ranked by score
optional $match
(specialty filter)
```

At ingest time, each note's chief complaint and clinical text are concatenated and embedded into a 768-dimensional vector using `nomic-embed-text`. At search time, the clinician's query is embedded with the same model and DocumentDB returns the most semantically similar cases — even when different clinical terminology is used.

## Open-source stack

| Component | Tool |
|---|---|
| Embedding model | [Ollama](https://ollama.com) `nomic-embed-text` (768 dimensions, runs locally) |
| Vector database | [DocumentDB OSS](https://github.com/microsoft/documentdb) via Docker |
| MongoDB driver | [PyMongo](https://pymongo.readthedocs.io/) |
| Web framework | [Flask](https://flask.palletsprojects.com/) |
| Language | Python 3.10+ |

## Prerequisites

- **Python 3.10+** — [python.org](https://python.org)
- **Docker Desktop** — [docker.com/products/docker-desktop](https://www.docker.com/products/docker-desktop)
- **Ollama** — [ollama.com/download](https://ollama.com/download)

After installing Ollama, pull the embedding model:

```bash
ollama pull nomic-embed-text
```

## Setup

### 1. Start DocumentDB OSS

**macOS / Linux / Git Bash:**
```bash
docker run -dt \
-p 10260:10260 \
-e USERNAME=docdbuser \
-e PASSWORD=Admin100! \
ghcr.io/microsoft/documentdb/documentdb-local:latest
```

**Windows PowerShell:**
```powershell
docker run -dt `
-p 10260:10260 `
-e USERNAME=docdbuser `
-e PASSWORD=Admin100! `
ghcr.io/microsoft/documentdb/documentdb-local:latest
```

### 2. Install dependencies

```bash
cd clinical-note-similarity-py
pip install -r requirements.txt
```

### 3. Configure environment variables

```bash
cp .env.example .env
```

Edit `.env` with your DocumentDB credentials:

| Variable | Default | Description |
|---|---|---|
| `DOCUMENTDB_URI` | — | Full MongoDB connection string |
| `DOCUMENTDB_DATABASE` | `clinicaldb` | Database name |
| `DOCUMENTDB_COLLECTION` | `notes` | Collection name |
| `OLLAMA_BASE_URL` | `http://127.0.0.1:11434` | Ollama server URL |
| `OLLAMA_EMBEDDING_MODEL` | `nomic-embed-text` | Embedding model |
| `FLASK_PORT` | `5001` | Port for the Flask web app |
| `NEAREST_NEIGHBORS` | `5` | Default number of results |
| `EMBEDDING_DIMENSIONS` | `768` | Must match the embedding model |

### 4. Upload clinical notes

Embeds all 20 sample notes and stores them in DocumentDB with a vector index:

```bash
python upload_notes.py
```

Expected output:
```
Loaded 20 clinical notes

Cleared existing collection

Generating embeddings and uploading...
[ 1/20] [Cardiology ] Acute ST-Elevation Myocardial Infarction...
[ 2/20] [Cardiology ] Unstable Angina...
...
[20/20] [Gastroenterology ] Crohn's Disease Flare...

Inserted 20 note documents
Vector index created (vector-ivf, dimensions: 768, similarity: COS)
```

### 5. Start the web app

```bash
python app.py
```

Open your browser at `http://localhost:5001`.

## Sample queries to try

| Query | Expected specialty |
|---|---|
| `Sudden severe chest pain with ST elevation, diaphoresis, left arm radiation` | Cardiology |
| `Wheezing, dyspnea not responding to albuterol inhaler` | Pulmonology |
| `Sudden onset facial droop, arm weakness, and speech difficulty` | Neurology |
| `Abdominal pain migrating to right lower quadrant with fever` | Gastroenterology |
| `Knee pop after pivoting with immediate swelling and instability` | Orthopedics |
| `Fatigue, weight gain, cold intolerance, hair thinning, and constipation` | Endocrinology |
| `Spreading skin redness with warmth and fever after minor skin break` | Dermatology |

## Document schema

Each clinical note document stored in DocumentDB:

| Field | Type | Description |
|---|---|---|
| `note_id` | string | Unique identifier (e.g. `CN001`) |
| `specialty` | string | Medical specialty |
| `diagnosis` | string | Primary diagnosis |
| `age_group` | string | `18-35`, `36-50`, `51-65`, `65+` |
| `sex` | string | `M` or `F` |
| `chief_complaint` | string | Presenting complaint in one sentence |
| `clinical_note` | string | De-identified clinical summary (3-5 sentences) |
| `icd_code` | string | ICD-10 diagnosis code |
| `outcome` | string | `admitted`, `discharged`, `referred`, `follow-up` |
| `embedding` | array | 768-dimensional vector (excluded from search results) |

## Specialties covered

The sample dataset includes 20 notes across 7 specialties:

| Specialty | Notes | Diagnoses |
|---|---|---|
| Cardiology | 4 | STEMI, Unstable Angina, Atrial Fibrillation, Heart Failure |
| Pulmonology | 3 | Pneumonia, Asthma Exacerbation, COPD Exacerbation |
| Neurology | 3 | Migraine, Ischemic Stroke, Seizure |
| Gastroenterology | 3 | Appendicitis, GERD, Crohn's Disease |
| Orthopedics | 3 | Distal Radius Fracture, Lumbar Disc Herniation, ACL Tear |
| Endocrinology | 2 | Type 2 Diabetes, Hashimoto's Thyroiditis |
| Dermatology | 2 | Cellulitis, Plaque Psoriasis |

## Project structure

```
clinical-note-similarity-py/
├── data/
│ └── clinical_notes.json # 20 fictional de-identified clinical notes
├── utils/
│ ├── __init__.py
│ ├── db.py # MongoDB client factory
│ └── embeddings.py # Ollama embedding helper
├── templates/
│ ├── index.html # Search page
│ └── note.html # Full note detail page
├── static/
│ └── style.css # Styles
├── upload_notes.py # Seeds DocumentDB with embeddings + vector index
├── cleanup.py # Drops the notes collection
├── app.py # Flask web application
├── requirements.txt
├── .env.example # Template — copy to .env
├── .gitignore
└── README.md
```

## Cleanup

Drop the notes collection when you are done:

```bash
python cleanup.py
```

## Disclaimer

This sample uses entirely fictional, de-identified clinical notes generated for demonstration purposes. It is not a medical device, clinical decision support system, or suitable for use with real patient data. Always consult qualified healthcare professionals for medical decisions.
108 changes: 108 additions & 0 deletions clinical-note-similarity-py/app.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
import os

from dotenv import load_dotenv

load_dotenv()

from flask import Flask, render_template, request
from utils.db import get_client, get_collection
from utils.embeddings import get_embedding

app = Flask(__name__)

_client = None
_col = None


def get_col():
global _client, _col
if _col is None:
_client = get_client()
_col = get_collection(_client)
return _col


Comment thread
khelanmodi marked this conversation as resolved.
def similarity_search(query: str, specialty: str, num_results: int) -> list:
col = get_col()
k = num_results if specialty == "all" else num_results * 4

embedding = get_embedding(query)

pipeline = [
{
"$search": {
"cosmosSearch": {
"vector": embedding,
"path": "embedding",
"k": k,
},
"returnStoredSource": True,
}
},
{
"$addFields": {"similarityScore": {"$meta": "searchScore"}}
},
{
"$project": {"embedding": 0}
},
]

if specialty and specialty != "all":
pipeline.append({"$match": {"specialty": specialty}})

pipeline.append({"$limit": num_results})

return list(col.aggregate(pipeline))


def get_specialties() -> list:
col = get_col()
return sorted(col.distinct("specialty"))


@app.route("/")
def index():
specialties = get_specialties()
return render_template("index.html", specialties=specialties)


@app.route("/search", methods=["POST"])
def search():
query = request.form.get("query", "").strip()
specialty = request.form.get("specialty", "all")
num_results = int(request.form.get("num_results", 5))
specialties = get_specialties()
Comment on lines +80 to +83

Comment on lines +80 to +84

Copilot AI Mar 19, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

num_results = int(request.form.get('num_results', 5)) can raise ValueError if the request is tampered with (or the field is missing/empty), causing a 500. Consider parsing num_results with a try/except and clamping to a small set of allowed values (or at least >= 1) similar to the handling in content-semantic-search-py/app.py.

Copilot uses AI. Check for mistakes.
results = []
error = None

if query:
try:
results = similarity_search(query, specialty, num_results)
except Exception as e:
error = str(e)
Comment thread
khelanmodi marked this conversation as resolved.
Outdated

return render_template(
"index.html",
query=query,
results=results,
specialty=specialty,
num_results=num_results,
specialties=specialties,
error=error,
)


@app.route("/note/<note_id>")
def note_detail(note_id):
col = get_col()
doc = col.find_one({"note_id": note_id}, {"embedding": 0})
if not doc:
return "Note not found", 404
return render_template("note.html", note=doc)


if __name__ == "__main__":
port = int(os.getenv("FLASK_PORT", 5001))
print(f"Starting Clinical Note Similarity Explorer on http://localhost:{port}")
app.run(debug=True, port=port)
Comment on lines +117 to +118
17 changes: 17 additions & 0 deletions clinical-note-similarity-py/cleanup.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
from dotenv import load_dotenv

load_dotenv()

from utils.db import get_client, get_collection


def main():
client = get_client()
col = get_collection(client)
col.drop()
print(f"Dropped collection: {col.full_name}")
client.close()


if __name__ == "__main__":
main()
Loading
Loading