dynamicsystemsarchitecture.org

Governance Layer

The real governance for the state-estimation pipeline — score aggregation, disagreement detection, a circuit breaker. Found with a real operator-precedence bug, fixed, and demonstrated with a concrete before/after case, not just described.

PurposeApply structural rules across multiple estimator outputs — confidence gating, disagreement detection, a circuit breaker on extreme+uncertain results — before producing one governed decision.
StatusVerified — bug found on inspection, fixed, and confirmed correct against a real execution on the Pipeline Orchestrator page
Built fromDirect code review, not assumed correct
Depends on
Superseded by
EvidenceConcrete before/after case below; real run on the Pipeline Orchestrator page.

A real bug, with a concrete case showing it actually mattered

The circuit breaker was written as:

if final_score < 0.2 or final_score > 0.8 and disagreement > 0.3:

Python binds and tighter than or, so this actually means final_score < 0.2 or (final_score > 0.8 and disagreement > 0.3) — a low score alone, with no disagreement required, always tripped the breaker. A high score only tripped it alongside high disagreement. Not symmetric, and not what the code's own comment ("extreme score with high disagreement") describes.

Concrete case, run, not just reasoned about:

final_score = 0.15   # low, but disagreement is also LOW -- estimators actually agree
disagreement = 0.05

Buggy original condition fires: True   <- WRONG: fires on low score alone, ignoring disagreement entirely
Fixed condition fires:          False  <- CORRECT: does not fire, disagreement too low to distrust the score

A real case where the estimators agreed with each other, and the original code would have raised a governance flag anyway — treating agreement as if it were the extreme-and-uncertain case the breaker was actually built to catch. Fixed by parenthesizing both sides symmetrically: (final_score < 0.2 or final_score > 0.8) and disagreement > 0.3.

Source (fixed)

#!/usr/bin/env python3
"""
governance_layer.py
NCFCA Governance Layer

Sits above individual estimators and applies higher-level rules:
- Invariant checking across estimators
- Conflict / disagreement detection
- Score aggregation with governance weighting
- Circuit breaker at pipeline level
- Audit logging of decisions
"""

from __future__ import annotations
from dataclasses import dataclass, field
from typing import Dict, List, Any, Optional
import numpy as np


@dataclass
class GovernanceResult:
    final_score: float
    confidence: float
    active_estimators: List[str]
    violations: List[str] = field(default_factory=list)
    metadata: Dict[str, Any] = field(default_factory=dict)


class GovernanceLayer:
    def __init__(self, min_confidence: float = 0.4, disagreement_threshold: float = 0.25):
        self.min_confidence = min_confidence
        self.disagreement_threshold = disagreement_threshold
        self.decision_log: List[str] = []

    def _log(self, message: str):
        self.decision_log.append(message)
        print(f"[GOVERNANCE] {message}")

    def govern(self, estimator_results: Dict[str, Dict[str, Any]]) -> GovernanceResult:
        if not estimator_results:
            return GovernanceResult(final_score=0.5, confidence=0.0, active_estimators=[])

        scores, confidences, active, violations = [], [], [], []

        for name, result in estimator_results.items():
            score = result.get("score", 0.5)
            conf = result.get("confidence", result.get("uncertainty", 0.5))
            if conf < self.min_confidence:
                violations.append(f"{name}: Low confidence ({conf:.2f})")
                self._log(f"Estimator {name} rejected due to low confidence")
                continue
            scores.append(score)
            confidences.append(conf)
            active.append(name)

        if not scores:
            self._log("All estimators rejected by governance")
            return GovernanceResult(final_score=0.5, confidence=0.0, active_estimators=[], violations=violations)

        disagreement = float(np.std(scores))
        if disagreement > self.disagreement_threshold:
            violations.append(f"High disagreement between estimators (std={disagreement:.3f})")
            self._log(f"High disagreement detected: {disagreement:.3f}")

        weights = np.array(confidences)
        weights = weights / np.sum(weights)
        final_score = float(np.average(scores, weights=weights))
        avg_confidence = float(np.mean(confidences))

        # FIXED: parenthesized symmetrically -- see the concrete case above
        if (final_score < 0.2 or final_score > 0.8) and disagreement > 0.3:
            violations.append("Extreme score with high disagreement — governance flag raised")
            self._log("Governance flag raised on extreme + disagreeing result")

        return GovernanceResult(
            final_score=final_score, confidence=avg_confidence, active_estimators=active,
            violations=violations,
            metadata={"num_estimators_run": len(estimator_results), "num_estimators_accepted": len(active), "disagreement": disagreement}
        )

Resolved: what "governance" actually means across this lab

A real investigation, worth documenting properly rather than just referencing. This module (GovernanceLayer) governs the state-estimation pipeline above — real estimator scores, real disagreement, real circuit breaking. A separate file, governance_v2.py (TestingGovernance), governs something else entirely: fault-injection testing for the execution kernel (ncfca_execution_kernel_v4.py) — its "decisions" are a seeded random draw (70% accept / 10% retry / 10% quarantine / 10% reject) used to stress-test the execution kernel's handling of corrupted or timed-out checkpoints, with two hardcoded overrides for deliberately injected faults.

These are not competing implementations of one concept — they're two real governance mechanisms for two different pipelines, and neither supersedes the other. Which one a new pipeline should use is itself a real, still-open selection question, tracked on the Master Kernel Library.

Comments

We'll only use this to contact you if we reply.

Comments

We'll only use this to contact you if we reply.