Skip to content
Merged
Show file tree
Hide file tree
Changes from 8 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang
- `prompt_source` identity gate: instruction and prompt boilerplate is not unique latent content and is not erased by a stopword list; `identity_recovery_rate` reports exact kind matches, with a contract test comparing correct recovery with an all-unique collapse on a mixed known-truth fixture (ADR 0004/0012).
- `location_membership` identity gate: geographic and market assignments are time-varying memberships, not permanent entity identity and not language channels; recovered location kinds match known truth at a higher computed rate than collapsing every assignment to entity identity (ADR 0003).
- `membership_target` identity gate: language, episode, template, department, and opportunity-pool memberships cannot collapse into the entity/project pair stored by migration `0006`; comparison-contract tests record recovered target kinds against an entity-collapse baseline (ADR 0003).

- `corpus_split` inferential-weight gate: only group-normalized ESS and uniform observation weights may enter an estimator; TF-IDF, BM25, and default global stopword deletion fail closed, with computed RMSE showing the retrieval surrogate recovers known shares worse than `group_normalized_ess`.
- `tepp_api` naruon live loopback HTTP/1.1 listener: `serve_one` installs a read/write deadline, requires a loopback `Host`, refuses `Transfer-Encoding` and NIM/proxy credential headers, parses `knowledge_cutoff` as RFC 3339 and refuses a future cutoff, keys analysis-run idempotency by tenant plus key, and proves both analysis-run and export POSTs over a real `TcpStream`. Not a production TLS/`$PORT` service (ADR 0011).
- `tepp_api` adaptive orchestration router (ADR 0010): versioned `direct`/`verify`/`committee`/`conductor`/`abstain` selection from CPU `f64` risk, ambiguity, evidence, and token-budget inputs; recorded stages, recursion, decomposition, access lists, and role-specific reasoning effort; fail-closed document-controlled policy/access/credentials; LLM plans remain proposals under deterministic statistical authority; comparable-budget ablation requires a direct baseline; credential-free contextual-orchestrator binding. Live NIM HTTP remains accepted-target.
- `tepp_api` purpose-bound provider-payload minimization: time-bounded `PurposeGrant` evaluation, fail-closed expired/not-yet-valid/inverted/cross-tenant/impossible-calendar denial, semantic UTC calendar validation, refusal to copy identity mappings into model-provider payloads or ordinary logs, preservation of opaque analytical identifiers and membership roles (no blanket PII mask), a separately authorized scientific re-identification path, and an internally bound FIPS 180-4 SHA-256 audit digest appended through `ReidentificationAuditSink` before disclosure.
Expand Down
1 change: 1 addition & 0 deletions DOCUMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ TEPP's approved PRD v0.4 and implementation plan are the primary product baselin
| Hourly NIM product-development operations | [`docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md`](docs/operations/HOURLY_NIM_PRODUCT_DEVELOPMENT.md) |
| Actions workflow fleet audit | [`docs/operations/ACTIONS_WORKFLOW_FLEET.md`](docs/operations/ACTIONS_WORKFLOW_FLEET.md) |
| Actions fleet research doctoring | [`docs/research/actions-workflow-fleet.md`](docs/research/actions-workflow-fleet.md) |
| Inferential TF-IDF/BM25/stopword refusal doctoring | [`docs/research/inferential-retrieval-weight-gate.md`](docs/research/inferential-retrieval-weight-gate.md) |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📝 Info: Duplicated documentation map diverges after one-sided edit

The file contains the whole documentation map twice. The new doctoring row is added only to the first copy; the second copy near line 80 is left without it, so the two tables now diverge. The duplication itself is pre-existing.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

| Mention-confidence Brier doctoring | [`docs/research/mention-confidence-brier.md`](docs/research/mention-confidence-brier.md) |
| Event-intelligence status-gate doctoring | [`docs/research/event-intelligence-status-gates.md`](docs/research/event-intelligence-status-gates.md) |
| VRAM budget / GPU fallback doctoring | [`docs/research/vram-budget-types.md`](docs/research/vram-budget-types.md) |
Expand Down
14 changes: 14 additions & 0 deletions crates/corpus_split/src/error.rs
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,10 @@ pub enum CorpusSplitError {
InvalidSplitConfiguration,
/// A document body was empty and has no Unicode identity.
EmptyCanonicalText,
/// A retrieval ranking score was treated as an inferential estimator weight.
InferentialRetrievalWeight,
/// Global stopword deletion was proposed as the default preprocessing rule.
DefaultStopwordDeletion,
}

impl fmt::Display for CorpusSplitError {
Expand All @@ -26,6 +30,8 @@ impl fmt::Display for CorpusSplitError {
Self::DuplicateDocumentIdentity => "duplicate document identity",
Self::InvalidSplitConfiguration => "invalid split configuration",
Self::EmptyCanonicalText => "empty canonical text",
Self::InferentialRetrievalWeight => "retrieval score is not an inferential weight",
Self::DefaultStopwordDeletion => "global stopword deletion is not the default rule",
};
formatter.write_str(message)
}
Expand Down Expand Up @@ -59,5 +65,13 @@ mod tests {
CorpusSplitError::EmptyCanonicalText.to_string(),
"empty canonical text"
);
assert_eq!(
CorpusSplitError::InferentialRetrievalWeight.to_string(),
"retrieval score is not an inferential weight"
);
assert_eq!(
CorpusSplitError::DefaultStopwordDeletion.to_string(),
"global stopword deletion is not the default rule"
);
}
}
144 changes: 144 additions & 0 deletions crates/corpus_split/src/inferential_weight.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
//! Retrieval scores and stopword deletion are not inferential split weights.

use crate::CorpusSplitError;

/// Proposed document or term scoring identity for a split or estimator input.
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
#[non_exhaustive]
pub enum WeightingScheme {
/// Kish / group-normalized observation weights.
GroupNormalizedEss,
/// Uniform observation weights.
Uniform,
/// TF-IDF retrieval ranking score.
TfIdf,
/// BM25 retrieval ranking score.
Bm25,
}

impl WeightingScheme {
/// Return whether this scheme may enter a statistical estimator as a weight.
#[must_use]
pub const fn is_inferential_weight(self) -> bool {
matches!(self, Self::GroupNormalizedEss | Self::Uniform)
}

/// Return the stable wire name.
#[must_use]
pub const fn wire_name(self) -> &'static str {
match self {
Self::GroupNormalizedEss => "group_normalized_ess",
Self::Uniform => "uniform",
Self::TfIdf => "tf_idf",
Self::Bm25 => "bm25",
}
}
}

/// Proposed token-deletion rule applied before estimation.
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
#[non_exhaustive]
pub enum TokenDeletionRule {
/// Keep tokens and model template, section, copied, and style as method structure.
PreserveAndModelBackground,
/// Delete tokens that appear on a global stopword list.
GlobalStopwordList,
}

impl TokenDeletionRule {
/// Return whether this rule is allowed as the default preprocessing policy.
#[must_use]
pub const fn is_default_allowed(self) -> bool {
matches!(self, Self::PreserveAndModelBackground)
}

/// Return the stable wire name.
#[must_use]
pub const fn wire_name(self) -> &'static str {
match self {
Self::PreserveAndModelBackground => "preserve_and_model_background",
Self::GlobalStopwordList => "global_stopword_list",
}
}
}

/// Refuse TF-IDF and BM25 as inferential estimator weights.
///
/// # Errors
///
/// Returns [`CorpusSplitError::InferentialRetrievalWeight`] unless `scheme` is
/// [`WeightingScheme::GroupNormalizedEss`] or [`WeightingScheme::Uniform`].
pub fn refuse_inferential_retrieval_weight(
scheme: WeightingScheme,
) -> Result<(), CorpusSplitError> {
if scheme.is_inferential_weight() {
Ok(())
} else {
Err(CorpusSplitError::InferentialRetrievalWeight)
}
}

/// Refuse global stopword deletion as the default preprocessing rule.
///
/// # Errors
///
/// Returns [`CorpusSplitError::DefaultStopwordDeletion`] unless `rule` is
/// [`TokenDeletionRule::PreserveAndModelBackground`].
pub fn refuse_default_stopword_deletion(rule: TokenDeletionRule) -> Result<(), CorpusSplitError> {
if rule.is_default_allowed() {
Ok(())
} else {
Err(CorpusSplitError::DefaultStopwordDeletion)
}
}

#[cfg(test)]
mod tests {
use super::{
TokenDeletionRule, WeightingScheme, refuse_default_stopword_deletion,
refuse_inferential_retrieval_weight,
};
use crate::CorpusSplitError;

#[test]
fn predicates_export_stable_wire_names_and_gates() {
assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight());
assert!(WeightingScheme::Uniform.is_inferential_weight());
assert!(!WeightingScheme::TfIdf.is_inferential_weight());
assert!(!WeightingScheme::Bm25.is_inferential_weight());
assert_eq!(
WeightingScheme::GroupNormalizedEss.wire_name(),
"group_normalized_ess"
);
assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform");
assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf");
assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25");
refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess");
refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform");
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::Bm25),
Err(CorpusSplitError::InferentialRetrievalWeight)
);

assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed());
assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed());
assert_eq!(
TokenDeletionRule::PreserveAndModelBackground.wire_name(),
"preserve_and_model_background"
);
assert_eq!(
TokenDeletionRule::GlobalStopwordList.wire_name(),
"global_stopword_list"
);
refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground)
.expect("preserve");
assert_eq!(
refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList),
Err(CorpusSplitError::DefaultStopwordDeletion)
);
}
}
9 changes: 9 additions & 0 deletions crates/corpus_split/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
mod connected_group;
mod document;
mod error;
mod inferential_weight;
mod rolling_origin;
mod snapshot;
mod unicode_identity;
Expand Down Expand Up @@ -37,6 +38,14 @@ pub use connected_group::build_connected_groups;
pub use document::CorpusDocument;
/// Fail-closed corpus-split errors.
pub use error::CorpusSplitError;
/// Token-deletion rule that may be proposed before estimation.
pub use inferential_weight::TokenDeletionRule;
/// Document or term scoring identity proposed as an estimator input.
pub use inferential_weight::WeightingScheme;
/// Refuse global stopword deletion as the default preprocessing rule.
pub use inferential_weight::refuse_default_stopword_deletion;
/// Refuse TF-IDF and BM25 as inferential estimator weights.
pub use inferential_weight::refuse_inferential_retrieval_weight;
/// Rolling-origin train/test window.
pub use rolling_origin::RollingOriginWindow;
/// Build ordered rolling-origin windows.
Expand Down
160 changes: 160 additions & 0 deletions crates/corpus_split/tests/inferential_weight_contract.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
//! TF-IDF, BM25, and global stopword deletion are not inferential inputs.

use corpus_split::{
CorpusSplitError, LeakageLink, LeakageLinkKind, TokenDeletionRule, WeightingScheme,
build_connected_groups, group_normalized_weights, refuse_default_stopword_deletion,
refuse_inferential_retrieval_weight,
};
use std::collections::BTreeMap;
use uuid::Uuid;

fn computed_rmse(truth: &[f64], recovered: &[f64]) -> f64 {
assert_eq!(truth.len(), recovered.len());
let n = f64::from(u32::try_from(truth.len()).expect("tiny fixture"));
let sse: f64 = truth
.iter()
.zip(recovered)
.map(|(truth_value, recovered_value)| {
let residual = truth_value - recovered_value;
residual * residual
})
.sum();
(sse / n).sqrt()
}

fn l1_normalize(values: &[f64]) -> Vec<f64> {
let total: f64 = values.iter().sum();
assert!(total > 0.0);
values.iter().map(|value| value / total).collect()
}

/// Classic summed TF-IDF retrieval scores used only as a negative surrogate.
fn tf_idf_document_scores(documents: &[&[&str]]) -> Vec<f64> {
let document_count = f64::from(u32::try_from(documents.len()).expect("tiny fixture"));
let mut document_frequency = BTreeMap::<&str, f64>::new();
for document in documents {
let mut seen = std::collections::BTreeSet::new();
for token in *document {
if seen.insert(*token) {
*document_frequency.entry(*token).or_insert(0.0) += 1.0;
}
}
}
documents
.iter()
.map(|document| {
let mut term_frequency = BTreeMap::<&str, f64>::new();
for token in *document {
*term_frequency.entry(*token).or_insert(0.0) += 1.0;
}
term_frequency
.into_iter()
.map(|(token, frequency)| {
let df = document_frequency.get(token).copied().unwrap_or(0.0);
frequency * (document_count / df).ln()
})
.sum()
})
.collect()
}

#[test]
fn allowed_observation_weights_pass_and_retrieval_scores_fail_closed() {
refuse_inferential_retrieval_weight(WeightingScheme::GroupNormalizedEss).expect("ess");
refuse_inferential_retrieval_weight(WeightingScheme::Uniform).expect("uniform");
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::Bm25),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
}

#[test]
fn global_stopword_deletion_is_not_the_default_rule() {
refuse_default_stopword_deletion(TokenDeletionRule::PreserveAndModelBackground)
.expect("preserve");
assert_eq!(
refuse_default_stopword_deletion(TokenDeletionRule::GlobalStopwordList),
Err(CorpusSplitError::DefaultStopwordDeletion)
);
}

#[test]
fn group_normalized_mass_recovers_true_shares_with_lower_rmse_than_tfidf() {
let truth = [0.40_f64, 0.10, 0.30, 0.20];
// Independent synthetic observation counts; not a scalar of `truth`.
let observation_mass = [41.0_f64, 9.0, 32.0, 18.0];
let documents: [&[&str]; 4] = [
&["report", "report", "report", "event"],
&["report", "unique"],
&["report", "event", "event"],
&["report", "report", "unique", "event"],
];
let document_ids: Vec<Uuid> = (0..truth.len()).map(|_| Uuid::now_v7()).collect();
let links: Vec<LeakageLink> = document_ids
.windows(2)
.map(|pair| LeakageLink {
left: pair[0],
right: pair[1],
kind: LeakageLinkKind::SameEpisode,
})
.collect();
let groups = build_connected_groups(&document_ids, &links);
let normalized_by_id: BTreeMap<Uuid, f64> = group_normalized_weights(
&groups,
&document_ids
.iter()
.copied()
.zip(observation_mass)
.collect::<Vec<_>>(),
)
.into_iter()
.collect();
let ess_recovered: Vec<f64> = document_ids
.iter()
.map(|document_id| *normalized_by_id.get(document_id).expect("normalized mass"))
.collect();
let tfidf_recovered = l1_normalize(&tf_idf_document_scores(&documents));
let ess_rmse = computed_rmse(&truth, &ess_recovered);
let tfidf_rmse = computed_rmse(&truth, &tfidf_recovered);
assert!(
ess_rmse < 0.05,
"independent observation mass must recover true shares; RMSE {ess_rmse}"
Comment thread
seonghobae marked this conversation as resolved.
);
assert!(
ess_rmse < tfidf_rmse,
"computed ESS RMSE {ess_rmse} must be below TF-IDF surrogate RMSE {tfidf_rmse}"
);
Comment thread
coderabbitai[bot] marked this conversation as resolved.
assert_eq!(
refuse_inferential_retrieval_weight(WeightingScheme::TfIdf),
Err(CorpusSplitError::InferentialRetrievalWeight)
);
}

#[test]
fn wire_names_and_predicates_are_stable() {
assert_eq!(
WeightingScheme::GroupNormalizedEss.wire_name(),
"group_normalized_ess"
);
assert_eq!(WeightingScheme::Uniform.wire_name(), "uniform");
assert_eq!(WeightingScheme::TfIdf.wire_name(), "tf_idf");
assert_eq!(WeightingScheme::Bm25.wire_name(), "bm25");
assert!(WeightingScheme::GroupNormalizedEss.is_inferential_weight());
assert!(WeightingScheme::Uniform.is_inferential_weight());
assert!(!WeightingScheme::TfIdf.is_inferential_weight());
assert!(!WeightingScheme::Bm25.is_inferential_weight());
assert_eq!(
TokenDeletionRule::PreserveAndModelBackground.wire_name(),
"preserve_and_model_background"
);
assert_eq!(
TokenDeletionRule::GlobalStopwordList.wire_name(),
"global_stopword_list"
);
assert!(TokenDeletionRule::PreserveAndModelBackground.is_default_allowed());
assert!(!TokenDeletionRule::GlobalStopwordList.is_default_allowed());
}
2 changes: 2 additions & 0 deletions docs/TRACEABILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,8 @@ The full APA 7th standards/literature register remains `docs/research/standards-
| multilingual shared latent semantic space | PRD; ADR 0004; ADR 0020 | `semantic_core` span-grounded units (active-PR); concept dictionary and shared latent estimator remaining | active-PR |
| TRSL-TM temporal/relational topic posterior and backend compatibility | ADR 0012; ADR 0004 | future `topic_measurement` | accepted-target |
| global P0 topic identity with activity/dormancy/reactivation | ADR 0012 | future topic lineage/activity state | accepted-target |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `corpus_split` inferential-weight gate on the active PR; estimator-side method model remains future | active-PR |
| report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; `prompt_source` prompt-versus-unique-content identity implemented-main; estimator-side method model remains future | partial |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | future semantic/method-source model | accepted-target |
| no default stopword deletion / no TF-IDF-BM25 inferential weighting | ADR 0004/0012; PRD/TRD | `stopword_deletion` default-list refusal on the active PR; TF-IDF/BM25 inferential-weight refusal remains accepted-target | partial |
Comment on lines +60 to 63

@devin-ai-integration devin-ai-integration Bot Aug 25, 2026

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📝 Info: Stale traceability rows contradict new gate status

The added row marks the corpus_split inferential-weight gate active-PR, but nearby existing rows still say the TF-IDF/BM25 refusal is remaining/accepted-target. These older rows are now stale and contradictory.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

| report template/section/copied/style/modality method effects | ADR 0004/0012; PRD/TRD | simulation truth factors implemented; estimator-side method model remains future | partial |
Expand Down
Loading
Loading