# Beyond Cost-Quality: Privacy-Aware Routing for Local-to-Cloud LLM Escalation

Canonical: https://careinstitute.ai/research/privacy-routing/
Published: 2026-02-15
Status: working
Author: Michael Turon
Code licence: Apache 2.0 (https://github.com/CAREInst)

## In short

**Thesis:** Every query has a sensitivity profile. The right architecture classifies first, routes to an admissible model set second, and audits the routing decision for every call.

## Key points

- Every query has a sensitivity profile, so the choice is not cloud or local: each request needs its own route.
- **Classify** each request by how sensitive it is: health data, personal data, privileged, classified, or general.
- **Route** it only to models cleared for its label, on a device or in the cloud, then pick the best one on cost and speed.
- **Audit** every choice, so a reviewer can check what went where, and why.
- The full text is pending. The flows on this site are illustrative, not measured data.

## The five classes


- **Health data.** “Summarize this patient’s latest lab results.” · Sent to: On-device model · Held from: Frontier model
- **Personal data.** No sample on this site yet.
- **Privileged.** “Summarize this deposition.” · Sent to: On-premises model, Frontier model (split) · Held from: Frontier model (case facts)
- **Classified.** “Help with this signal pattern.” · Sent to: Accredited system on the secure network · Held from: Frontier model
- **General.** “Does Drug A clash with Drug B?” · Sent to: Frontier model · Held from: None

Samples: illustrative flow — not measured data.

Working paper. Full text pending.

## What we don’t know yet

- How much time a classifier model adds when it runs on the same device. We have not measured that yet.
