A complete, azd up-deployable sample that gives a prompt agent in Microsoft
Foundry Agent Service durable, user-scoped memory powered by Azure Cosmos DB
and Microsoft Agent Framework.
Does
azd updeploy everything? Yes. It provisions the Foundry account and project, chat and embedding model deployments, Azure Cosmos DB, and RBAC. A post-provision hook then creates the PromptAgent in Foundry Agent Service and runs a live two-conversation memory test. The only local component is the optional Chainlit browser UI that you start after deployment.
| Component | Created by | Purpose |
|---|---|---|
| Azure resource group | Azure Developer CLI | Contains the sample resources |
| Microsoft Foundry account and project | Bicep | Hosts models and Agent Service |
gpt-5-mini deployment |
Bicep | Generates agent responses and extracts memory |
text-embedding-3-large deployment |
Bicep | Embeds memory for semantic retrieval |
| Azure Cosmos DB for NoSQL account | Bicep | Stores turns and derived memory |
ai_memory database |
Bicep | Memory toolkit database |
| Vector and full-text search capabilities | Bicep | Hybrid memory retrieval |
| Foundry and Cosmos DB role assignments | Bicep | Keyless data-plane access for the deploying identity |
cosmos-memory-sample-agent PromptAgent |
Post-provision Python hook | The deployed Foundry Agent Service agent |
| Python virtual environment and dependencies | Post-provision hook | Runs the sample and validation |
| Cross-conversation memory test | Post-provision hook | Makes deployment fail if live recall is not demonstrated |
Cosmos DB local authentication is disabled. The code uses
DefaultAzureCredential and Microsoft Entra ID instead of account keys or
connection strings.
flowchart LR
Browser[Chainlit browser chat] --> Runtime[Microsoft Agent Framework]
Runtime --> Agent[Foundry Agent Service<br/>PromptAgent]
Agent --> Chat[Chat model deployment]
Runtime <--> Memory[CosmosMemoryContextProvider]
Memory --> Embedding[Embedding model deployment]
Memory <--> Cosmos[(Azure Cosmos DB<br/>turns, facts, summaries, profiles)]
The provider runs automatically around each agent call:
before_runretrieves relevant memory from Cosmos DB and adds it to the agent context.after_runstores the new turn and starts fact, summary, and profile extraction.
Conversation identity and memory identity are deliberately separate:
flowchart LR
A1[Theo<br/>Conversation 1] -->|save fact for user_id=theo| M[(Theo's Cosmos memory)]
A1 -->|new conversation| A2[Theo<br/>Conversation 2]
M -->|recall fact| A2
C[Casey<br/>user_id=casey] --> CM[(Casey's isolated memory)]
A new Agent Framework session starts a clean conversation. Keeping the same
user_id lets memory follow the user; changing it creates a separate retrieval
scope.
- Azure Developer CLI (
azd) - Azure CLI (
az) - Python 3.11 or later
- An Azure subscription with permission to create resources and assign roles
- Foundry model quota in the selected region
Because the template creates role assignments, use an identity with Owner or
User Access Administrator plus permission to create the resources. Model
availability and quota vary by region. This sample is validated in Sweden
Central, which offers both gpt-5-mini (GlobalStandard) and
text-embedding-3-large (regional Standard) with quota. The regional Standard
embedding SKU is not offered in every region, so prefer Sweden Central unless you
have confirmed both SKUs in your target region.
Clone your fork and sign in:
git clone https://github.com/TheovanKraay/foundry-cosmos-memory.git
cd foundry-cosmos-memory
az login
azd auth login
azd upazd up prompts for an environment name, subscription, and region. Provisioning
and the post-provision validation can take several minutes. Choose Sweden
Central at the region prompt, or set it up front with
azd env set AZURE_LOCATION swedencentral before running azd up.
A successful deployment ends with output similar to:
[Conversation 1] Teaching the agent a durable fact...
[Flush] Waiting for background memory extraction to persist...
[Conversation 2] New conversation, same user. Asking it to recall...
Recalled peanut allergy in a NEW conversation: True
PASS - Cosmos memory was injected into the Foundry agent run.
The test generates a new memory-test-<uuid> user every time. This prevents old
data from producing a false pass. If the second conversation does not mention the
newly taught peanut allergy, the script exits nonzero and azd up fails.
The cloud resources and agent are now deployed. The post-provision step installs
only the packages needed to create the agent and run the memory test, so install
the optional Chainlit UI dependencies once before starting the browser chat.
Export the azd outputs to a local .env file and start Chainlit.
azd env get-values | Set-Content .env
.\.venv\Scripts\python.exe -m pip install -r requirements-ui.txt --pre
.\.venv\Scripts\python.exe -m chainlit run src/chat.pyazd env get-values > .env
. .venv/bin/activate
python -m pip install -r requirements-ui.txt --pre
python -m chainlit run src/chat.pyOpen http://localhost:8000. The UI is local, but every
agent run, model call, and memory operation uses the Azure resources deployed by
azd up.
Try this sequence:
- On the first screen, choose the
theodemo user, or choose Type my own. - Send
Remember that my favorite color is vermilion. - Wait for Save long-term memory to finish.
- Start a New chat (top-left) and choose
theoagain. - Ask
What is my favorite color?The agent recalls it from Cosmos DB memory. - Start a New chat (top-left), choose
casey, then ask the same question to demonstrate isolation.
Do not enter personal, confidential, or sensitive information. Typed user IDs are only a simple demonstration mechanism; production applications should derive the memory identity from an authenticated principal.
The app also supports:
/new- create a fresh conversation for the current memory user/user <id>- switch memory users and create a fresh conversation/help- show the available commands
The .env exported above contains resource names and endpoints, not secrets.
.\.venv\Scripts\python.exe -m src.run_memory_test. .venv/bin/activate
python -m src.run_memory_testsrc/agent_runtime.py constructs the same runtime used by
the browser app and deterministic test:
memory = CosmosMemoryContextProvider(
cosmos_endpoint=config.COSMOS_ENDPOINT,
cosmos_database=config.COSMOS_DATABASE,
foundry_endpoint=config.FOUNDRY_PROJECT_ENDPOINT,
embedding_model=config.EMBEDDING_MODEL,
chat_model=config.CHAT_MODEL,
credential=credential,
memory_types=["fact", "procedural", "episodic"],
)
agent = FoundryAgent(
project_endpoint=config.FOUNDRY_PROJECT_ENDPOINT,
agent_name=config.FOUNDRY_AGENT_NAME,
agent_version=os.getenv("FOUNDRY_AGENT_VERSION"),
credential=credential,
context_providers=[memory],
)Each conversation gets a new session while the stable user ID is stored in the provider's own state:
session = agent.create_session()
session.state.setdefault(memory.source_id, {})["user_id"] = user_idTo skip provisioning, copy .env.example to .env, fill in the
existing Cosmos DB and Foundry project values, install requirements.txt, and run:
python -m src.create_agent
python -m src.run_memory_testPersist the FOUNDRY_AGENT_NAME and FOUNDRY_AGENT_VERSION printed by the first
command in .env before running the test.
Use a region that offers both gpt-5-mini (GlobalStandard) and
text-embedding-3-large (regional Standard) with quota. The regional Standard
embedding SKU is not available in every region (for example, West US, North
Central US, and South Central US); Sweden Central, East US, West US 3, Canada
East, and Australia East do offer it. The chat deployment uses GlobalStandard
with capacity 50 and the embedding deployment uses Standard with capacity 10.
You can adjust names, versions, SKU, and capacity in
infra/main.bicep.
The deploying identity must be allowed to create Azure role assignments. Use Owner or User Access Administrator at the subscription or target resource-group scope.
Role propagation can take several minutes. The template assigns Foundry User
at both account and project scope and Cognitive Services OpenAI User at the
account scope. Retry azd up after propagation completes.
Run it again, then inspect the ai_memory database in Cosmos DB Data Explorer.
Look for turn and extracted fact documents under the generated
memory-test-<uuid> identity. The test calls flush() before recall so background
extraction completes before the second conversation.
.
|-- azure.yaml # azd project and lifecycle hook
|-- infra/
| |-- main.bicep # Foundry, models, Cosmos DB, and RBAC
| `-- main.parameters.json
|-- hooks/
| |-- postprovision.ps1 # Windows setup, agent creation, live test
| `-- postprovision.sh # POSIX setup, agent creation, live test
|-- src/
| |-- agent_runtime.py # Shared agent, provider, and session lifecycle
| |-- chat.py # Chainlit browser experience
| |-- config.py # Environment-driven configuration
| |-- create_agent.py # Idempotent PromptAgent creation
| `-- run_memory_test.py # Fresh-user cross-conversation test
|-- .chainlit/config.toml
|-- .env.example
|-- requirements.txt # Core runtime (agent creation + memory test)
`-- requirements-ui.txt # Optional Chainlit browser UI
azd down --purge--purge also removes the soft-deleted Foundry account so its name is released.