fluxVendor-neutral control plane for heterogeneous agent fleets — one stitched OpenTelemetry trace across Quench (TS) and Anvil (Python), with per-agent cost and eval-health roster.
quenchCI/deployment incident-triage agent with human-gated writes — 14 trajectory evals, 78 checks, 12 real recorded GitHub Actions failures, dry-run by default.
anvilSupport agent over real FastAPI docs with retrieval and answer evals — hybrid retrieval, rerank, grounded citations, HITL issue tools, recall@5 0.567.
ingotToken-to-deliverable analytics for AI coding agents — joins usage telemetry with git history: spend, cache economics, cost per shipped commit.
assayEnterprise document extraction with per-field evals — real SROIE receipts, synthetic stress set, constrained decoding, validation, confidence routing.
crucibleForensic eval harness for self-hosted models, head of the Crucible family (with Crucible Lab and Bellows) — capability, refusal, tool-calling, RAG, base-vs-abliterated deltas, judge grading, CI gates.
bellowsThe Crucible family's MCP server for operating a local llama.cpp fleet — typed tool calls to scan GGUF models, supervise llama-server, smoke-test, and query Crucible history.