Skip to content
Merged
Show file tree
Hide file tree
Changes from 9 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang

### Added

- `evidence_core` embedded-image units: `data:image/<type>;base64,...` URIs keep their original source spans and media types, and cannot be used as lexical inference text.
- `tepp_api` naruon live loopback HTTP/1.1 listener: `serve_one` installs a read/write deadline, requires a loopback `Host`, refuses `Transfer-Encoding` and NIM/proxy credential headers, parses `knowledge_cutoff` as RFC 3339 and refuses a future cutoff, keys analysis-run idempotency by tenant plus key, and proves both analysis-run and export POSTs over a real `TcpStream`. Not a production TLS/`$PORT` service (ADR 0011).
- `tepp_api` adaptive orchestration router (ADR 0010): versioned `direct`/`verify`/`committee`/`conductor`/`abstain` selection from CPU `f64` risk, ambiguity, evidence, and token-budget inputs; recorded stages, recursion, decomposition, access lists, and role-specific reasoning effort; fail-closed document-controlled policy/access/credentials; LLM plans remain proposals under deterministic statistical authority; comparable-budget ablation requires a direct baseline; credential-free contextual-orchestrator binding. Live NIM HTTP remains accepted-target.
- `tepp_api` purpose-bound provider-payload minimization: time-bounded `PurposeGrant` evaluation, fail-closed expired/not-yet-valid/inverted/cross-tenant/impossible-calendar denial, semantic UTC calendar validation, refusal to copy identity mappings into model-provider payloads or ordinary logs, preservation of opaque analytical identifiers and membership roles (no blanket PII mask), a separately authorized scientific re-identification path, and an internally bound FIPS 180-4 SHA-256 audit digest appended through `ReidentificationAuditSink` before disclosure.
Expand Down
1 change: 1 addition & 0 deletions DOCUMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ TEPP's approved PRD v0.4 and implementation plan are the primary product baselin
| Hourly NIM product-development operations | [`docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md`](docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md) |
| Actions workflow fleet audit | [`docs/operations/ACTIONS_WORKFLOW_FLEET.md`](docs/operations/ACTIONS_WORKFLOW_FLEET.md) |
| Actions fleet research doctoring | [`docs/research/actions-workflow-fleet.md`](docs/research/actions-workflow-fleet.md) |
| Embedded-image unit doctoring | [`docs/research/embedded-image-units.md`](docs/research/embedded-image-units.md) |
| Retention/deletion/legal-hold doctoring | [`docs/research/retention-deletion-legal-hold.md`](docs/research/retention-deletion-legal-hold.md) |
| Provider-payload minimization doctoring | [`docs/research/provider-payload-minimization.md`](docs/research/provider-payload-minimization.md) |
| Adaptive orchestration router doctoring | [`docs/research/adaptive-orchestration-router.md`](docs/research/adaptive-orchestration-router.md) |
Expand Down
6 changes: 6 additions & 0 deletions crates/evidence_core/src/error.rs
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,10 @@ pub enum EvidenceError {
InvalidLayoutBounds,
/// Layout coordinates exceeded the enclosing page.
LayoutOutOfBounds,
/// A base64 image data URI was treated as lexical inference text.
EmbeddedImageIsNotLexicalText,
/// An embedded image data URI declared an implausible image media type.
ImplausibleImageMediaType,
}

impl fmt::Display for EvidenceError {
Expand All @@ -70,6 +74,8 @@ impl fmt::Display for EvidenceError {
Self::InvalidPageGeometry => "page geometry must be finite and positive",
Self::InvalidLayoutBounds => "layout bounds must be finite, nonnegative, and nonempty",
Self::LayoutOutOfBounds => "layout bounds exceed the page geometry",
Self::EmbeddedImageIsNotLexicalText => "embedded image is not lexical text",
Self::ImplausibleImageMediaType => "embedded image media type is implausible",
Comment thread
seonghobae marked this conversation as resolved.
};
formatter.write_str(message)
}
Expand Down
217 changes: 217 additions & 0 deletions crates/evidence_core/src/image_unit.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,217 @@
//! Embedded `data:image` units that keep their original source location.

use crate::{DocumentRecord, EvidenceError, SourceSpan};

const DATA_IMAGE_PREFIX: &str = "data:image/";
const BASE64_MARK: &str = ";base64,";

/// Image media types accepted as plausible by [`embedded_image_units`].
///
/// The set is deliberately conservative and tracks widely registered or
/// de facto standard image subtypes; anything else fails closed instead of
/// yielding a bogus embedded-image unit.
const PLAUSIBLE_IMAGE_MEDIA_TYPES: [&str; 14] = [
"image/apng",
"image/avif",
"image/bmp",
"image/gif",
"image/heic",
"image/heif",
"image/jpeg",
"image/jpg",
"image/png",
"image/svg+xml",
"image/tiff",
"image/vnd.microsoft.icon",
"image/webp",
"image/x-icon",
];

/// One embedded image located in a document body.
#[derive(Clone, Copy, Debug, PartialEq)]
pub struct EmbeddedImageUnit<'document> {
span: SourceSpan,
media_type: &'document str,
}

impl<'document> EmbeddedImageUnit<'document> {
/// Exact source span of the data URI, including the `data:image/` prefix.
#[must_use]
pub const fn span(self) -> SourceSpan {
self.span
}

/// Declared image media type (`image/png`, `image/jpeg`, …).
#[must_use]
pub const fn media_type(self) -> &'document str {
self.media_type
}
}

/// Locate `data:image/<type>;base64,...` units and retain their original spans.
///
/// Only plausible image media types are accepted: a candidate URI whose
/// declared media type is not in [`PLAUSIBLE_IMAGE_MEDIA_TYPES`] fails the
/// whole parse so malformed bodies cannot produce bogus units.
///
/// # Errors
///
/// Returns [`EvidenceError::EmptySourceSpan`] when the document contains no
/// well-formed embedded image URI, and
/// [`EvidenceError::ImplausibleImageMediaType`] when a candidate URI
/// declares an implausible image media type.
pub fn embedded_image_units(
document: &DocumentRecord,
) -> Result<Vec<EmbeddedImageUnit<'_>>, EvidenceError> {
let text = document.text();
let mut units = Vec::new();
let mut search_from = 0usize;
while let Some(relative) = text[search_from..].find(DATA_IMAGE_PREFIX) {
let start = search_from + relative;
let after_prefix = start + DATA_IMAGE_PREFIX.len();
let Some(mark_rel) = text[after_prefix..].find(BASE64_MARK) else {
search_from = after_prefix;
continue;
};
let media_end = after_prefix + mark_rel;
let payload_start = media_end + BASE64_MARK.len();
let payload_end = payload_start
+ text[payload_start..]
.find(|ch: char| !is_base64_payload_char(ch))
.unwrap_or(text.len() - payload_start);
if payload_end == payload_start {
search_from = payload_start;
continue;
}
let media_type = &text[start + "data:".len()..media_end];
Comment thread
devin-ai-integration[bot] marked this conversation as resolved.
if media_type.contains(DATA_IMAGE_PREFIX) {
search_from = after_prefix;
continue;
}
Comment thread
seonghobae marked this conversation as resolved.
if !is_plausible_image_media_type(media_type) {
return Err(EvidenceError::ImplausibleImageMediaType);
}
let scalar_start = text[..start].chars().count();
let scalar_end = scalar_start + text[start..payload_end].chars().count();
let span = SourceSpan::new(document, start, payload_end, scalar_start, scalar_end, None)?;
units.push(EmbeddedImageUnit { span, media_type });
Comment thread
seonghobae marked this conversation as resolved.
Outdated
search_from = payload_end;
}
if units.is_empty() {
return Err(EvidenceError::EmptySourceSpan);
Comment thread
seonghobae marked this conversation as resolved.
}
Ok(units)
}

/// Refuse using a document body that still contains an embedded image as
/// lexical inference text.
///
/// # Errors
///
/// Returns [`EvidenceError::InvalidWirePayload`] for empty input and
/// [`EvidenceError::EmbeddedImageIsNotLexicalText`] when a `data:image`
/// base64 URI is present.
pub fn refuse_base64_image_as_lexical_text(text: &str) -> Result<(), EvidenceError> {
if text.is_empty() {
return Err(EvidenceError::InvalidWirePayload);
}
if text.contains(DATA_IMAGE_PREFIX) && text.contains(BASE64_MARK) {
return Err(EvidenceError::EmbeddedImageIsNotLexicalText);
Comment thread
seonghobae marked this conversation as resolved.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated
}
Ok(())
}

fn is_base64_payload_char(ch: char) -> bool {
ch.is_ascii_alphanumeric() || matches!(ch, '+' | '/' | '=')
}

/// Report whether a declared media type is a plausible image media type.
fn is_plausible_image_media_type(media_type: &str) -> bool {
PLAUSIBLE_IMAGE_MEDIA_TYPES.contains(&media_type)
}
Comment thread
seonghobae marked this conversation as resolved.

#[cfg(test)]
mod tests {
use super::{embedded_image_units, refuse_base64_image_as_lexical_text};
use crate::{DocumentRecord, EvidenceError, SourceArtifact};

#[test]
fn jpeg_uri_and_incomplete_prefix_are_classified() {
let text = "x data:image/jpeg;base64,/9j/4AA= y data:image/gif y";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
let units = embedded_image_units(&document).expect("jpeg");
assert_eq!(units.len(), 1);
assert_eq!(units[0].media_type(), "image/jpeg");
refuse_base64_image_as_lexical_text("plain note").expect("plain");
refuse_base64_image_as_lexical_text("data:image/png").expect("incomplete image");
assert_eq!(
refuse_base64_image_as_lexical_text("data:image/png;base64,AAAA"),
Err(EvidenceError::EmbeddedImageIsNotLexicalText)
);

let empty_text = "data:image/png;base64, following text";
let empty_artifact = SourceArtifact::from_bytes(empty_text.as_bytes()).expect("artifact");
let empty_document =
DocumentRecord::from_text(empty_artifact.id(), empty_text).expect("document");
assert_eq!(
embedded_image_units(&empty_document),
Err(EvidenceError::EmptySourceSpan)
);
}

#[test]
fn implausible_media_types_fail_closed() {
let text = "data:image/not-a-type;base64,AAAA";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
assert_eq!(
embedded_image_units(&document),
Err(EvidenceError::ImplausibleImageMediaType)
);
assert_eq!(
refuse_base64_image_as_lexical_text(text),
Err(EvidenceError::EmbeddedImageIsNotLexicalText)
);
}

#[test]
fn common_raster_media_types_are_accepted() {
let text = "a data:image/png;base64,AAAA b data:image/jpeg;base64,BBBB \
c data:image/webp;base64,CCCC d data:image/gif;base64,DDDD e";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
let units = embedded_image_units(&document).expect("units");
let media_types: Vec<&str> = units.iter().map(|unit| unit.media_type()).collect();
assert_eq!(
media_types,
vec!["image/png", "image/jpeg", "image/webp", "image/gif"]
);
}

#[test]
fn empty_payload_is_not_an_image_unit() {
let text = "data:image/png;base64,";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
assert_eq!(
embedded_image_units(&document),
Err(EvidenceError::EmptySourceSpan)
);
}

#[test]
fn malformed_image_prefix_does_not_swallow_later_valid_image() {
let text = "data:image/gif then data:image/png;base64,AAAA";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");

let units = embedded_image_units(&document).expect("png");
assert_eq!(units.len(), 1);
assert_eq!(units[0].media_type(), "image/png");
assert_eq!(
units[0].span().byte_start(),
text.find("data:image/png").expect("png start")
);
}
}
10 changes: 9 additions & 1 deletion crates/evidence_core/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,15 @@
//! records, source spans whose byte, Unicode-scalar, page, and layout
//! coordinates are validated before entering later temporal or psychometric
//! layers, and strict versioned JSON wire contracts that reconstruct records
//! only through the same domain validation boundary.
//! only through the same domain validation boundary. Embedded `data:image`
//! units keep their original offsets and are not lexical inference text.

mod artifact;
mod digest;
mod document;
mod error;
mod identifier;
mod image_unit;
mod span;
mod wire;

Expand All @@ -27,6 +29,12 @@ pub use document::DocumentRecord;
pub use error::EvidenceError;
/// A validated RFC 9562 `UUIDv7` evidence identifier.
pub use identifier::EvidenceId;
/// One embedded image located in a document body.
pub use image_unit::EmbeddedImageUnit;
/// Locate `data:image` base64 units with exact source spans.
pub use image_unit::embedded_image_units;
/// Refuse treating an embedded image URI as lexical inference text.
pub use image_unit::refuse_base64_image_as_lexical_text;
/// A validated page-relative location for source evidence.
pub use span::PageLocation;
/// An exact byte, Unicode-scalar, and optional page/layout span.
Expand Down
75 changes: 75 additions & 0 deletions crates/evidence_core/tests/embedded_image_contract.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
//! Embedded base64 images keep their original location and are not lexical text.

use evidence_core::{
DocumentRecord, EvidenceError, SourceArtifact, embedded_image_units,
refuse_base64_image_as_lexical_text,
};

#[test]
fn data_uri_recovers_exact_span_and_media_type() {
let uri = "data:image/png;base64,iVBORw0KGgo=";
let text = format!("Before the figure.\n\n{uri}\n\nAfter the figure. data:image/gif y");
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), &text).expect("document");

let units = embedded_image_units(&document).expect("units");
assert_eq!(units.len(), 1);
assert_eq!(units[0].media_type(), "image/png");
assert_eq!(
&document.text()[units[0].span().byte_start()..units[0].span().byte_end()],
uri
);
assert_eq!(
refuse_base64_image_as_lexical_text(document.text()),
Err(EvidenceError::EmbeddedImageIsNotLexicalText)
);
refuse_base64_image_as_lexical_text("data:image/png").expect("incomplete image");
refuse_base64_image_as_lexical_text("Before the figure.").expect("plain text");
refuse_base64_image_as_lexical_text("data:image/gif y").expect("incomplete image marker");
}

#[test]
fn implausible_media_types_fail_closed_and_common_types_are_accepted() {
let malformed = "data:image/not-a-type;base64,AAAA";
let artifact = SourceArtifact::from_bytes(malformed.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), malformed).expect("document");
assert_eq!(
embedded_image_units(&document),
Err(EvidenceError::ImplausibleImageMediaType)
);

let text = "a data:image/png;base64,AAAA b data:image/jpeg;base64,BBBB \
c data:image/webp;base64,CCCC d data:image/gif;base64,DDDD e";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
let units = embedded_image_units(&document).expect("units");
let media_types: Vec<&str> = units.iter().map(|unit| unit.media_type()).collect();
assert_eq!(
media_types,
vec!["image/png", "image/jpeg", "image/webp", "image/gif"]
);
}

#[test]
fn documents_without_images_and_empty_payloads_fail_closed() {
let text = "No figures in this note.";
let artifact = SourceArtifact::from_bytes(text.as_bytes()).expect("artifact");
let document = DocumentRecord::from_text(artifact.id(), text).expect("document");
assert_eq!(
embedded_image_units(&document),
Err(EvidenceError::EmptySourceSpan)
);

assert_eq!(
refuse_base64_image_as_lexical_text(""),
Err(EvidenceError::InvalidWirePayload)
);
let empty_payload = "data:image/png;base64,";
let empty_artifact = SourceArtifact::from_bytes(empty_payload.as_bytes()).expect("artifact");
let empty_document =
DocumentRecord::from_text(empty_artifact.id(), empty_payload).expect("document");
assert_eq!(
embedded_image_units(&empty_document),
Err(EvidenceError::EmptySourceSpan)
);
}
4 changes: 4 additions & 0 deletions crates/evidence_core/tests/records_and_spans_contract.rs
Original file line number Diff line number Diff line change
Expand Up @@ -329,6 +329,10 @@ fn every_record_validation_error_has_a_stable_message() {
EvidenceError::LayoutOutOfBounds,
"layout bounds exceed the page geometry",
),
(
EvidenceError::EmbeddedImageIsNotLexicalText,
"embedded image is not lexical text",
),
];

for (error, expected) in cases {
Expand Down
1 change: 1 addition & 0 deletions docs/TRACEABILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ The full APA 7th standards/literature register remains `docs/research/standards-
| Requirement / decision | Canonical basis | Source/evidence boundary | Maturity |
|---|---|---|---|
| immutable source evidence and exact spans | PRD; Architecture; ADR 0008 | `evidence_core`, Task 2 tests/doctoring; `persistence_postgres` source-artifact SQL insert/lookup plus idempotent retry (#40 implemented-main) | implemented-main |
| embedded image location and non-lexical treatment | ADR 0008; research | `evidence_core` data-URI spans on the active PR | active-PR |
| Rust numerical authority / CPU `f64` reference | ADR 0001 | current workspace foundation; future estimators | partial |
| Rust workspace/quality foundation | ADR 0007 | workspace/CI/repository contract | implemented-main |
| six distinct clocks and uncertain intervals | PRD; ADR 0002 | PR #8 `temporal_core` on protected main; PR #5 historical only | implemented-main |
Expand Down
4 changes: 2 additions & 2 deletions docs/adr/0009-purpose-bound-pii-governance.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# ADR 0009 — Purpose-bound PII governance without blanket masking

**Decision status:** Accepted
**Decision status:** Accepted
**Implementation maturity:** partial — persistence retention/deletion/legal-hold (migration `0007`) is implemented-main; purpose-bound provider-payload minimization (expired-purpose denial, log/source separation, separately authorized re-identification) is on the active PR and is not implemented-main until exact-head checks, review, and protected-main integration complete; deployment/provider-region evidence remains accepted-target

**Date:** 2026-08-10
**Date:** 2026-08-10
Comment thread
seonghobae marked this conversation as resolved.
**Supersedes:** None.

## Context
Expand Down
Loading
Loading