Training data built from a real, running governance system
These datasets were assembled from SCBE-AETHERMOORE source, documentation, experiments, and project sessions. Counts and formats are concrete; benchmark results are project-authored and should be read with the scope and limitations in the evidence ledger.
Targeted datasets for specific use cases.
Each pack is documented and formatted for QLoRA-style pipelines. Validate its schema, license, split, and task fit with your chosen trainer before spending compute.
SCBE Governance SFT Pack
5,188 supervised fine-tuning pairs covering governance, system-level instructions, and codebase understanding.
- Merged multi-source SFT corpus (~5,188 pairs) covering governance + system + codebase
- QLoRA-ready JSONL with governance decisions, risk labels, and Sacred Tongue routing tags
- Sources: HF
scbe-aethermoore-training-data, internalsft_repo_merged.jsonl - Use as one input to governance-oriented fine-tuning experiments; results depend on the base model, mix, and evaluation design
Red Team Fortress
91 project-authored adversarial prompts across 10 categories, labeled by failure layer (L1-L14). The current 91/91 result is regression-corpus fit, not a generalization claim.
- 91 attack prompts across 10 adversarial categories
- Each prompt labeled by failure layer (L1 through L14)
- Compliance evals and attack scenario theory docs included
- Stress-test AI models and build adversarial training sets
- Also available as a free preview on HuggingFace (95 downloads)
Six Tongues Conlang + Tokenization Pack
Constructed language design docs, tokenization theory, and bijective encoding patterns from the Six Sacred Tongues system.
- theory_doc_conlang_intent.jsonl -- language design theory
- Tongues session transcripts
- Spiralverse codex SFT pairs
- Novel tokenization research and conlang AI training
- Creative AI fine-tuning for unique output patterns
Spiralverse Session Transcripts
48 session transcript files across 11 categories of structured RPG and worldbuilding dialogue.
- 48 session files: game, NPC roundtable, game design, lore
- Architecture, tongues, gacha, math, DM, music, space commerce
- ~400KB of structured dialogue data
- RPG AI training and game NPC dialogue
- Interactive fiction fine-tuning
Theory Documents Bundle
6 deep-dive theory documents covering the full intellectual foundation of the SCBE-AETHERMOORE framework.
- Architecture theory -- full pipeline design rationale
- Spiralverse lore -- worldbuilding knowledge injection
- Security attacks -- adversarial taxonomy and defenses
- GeoSeal crypto -- bijective encryption theory
- Conlang intent -- constructed language design principles
- Patent claims -- USPTO #63/961,403 methodology
The Full Arsenal.
Everything above in one download. Save $107 versus buying each pack individually.
The Full Arsenal
The combined SCBE-AETHERMOORE data collection for governed-AI experiments. A usable fine-tune still requires a compatible base model, trainer, compute budget, validation split, and independent evaluation.
- 5,188+ SFT training pairs
- 91 red team adversarial prompts
- 6 deep-dive theory documents
- 48 session transcripts (11 categories)
- Full conlang + tokenization pack
- Context capsules
- Knowledge base files
- Eval suites
Start here for free.
The framework is MIT-licensed. The HuggingFace datasets are free to download. No strings attached.
SCBE-AETHERMOORE
The experimental 14-layer pipeline, Sacred Tongues, and hyperbolic cost engine. Open source for inspection and testing; production readiness depends on your threat model and validation.
- Full 14-layer security pipeline
- Sacred Tongues tokenizers
- Hyperbolic cost scaling (H(d,R) = R^(d²))
- Project-authored evaluation artifacts with scope and limitations published separately
npm i scbe-aethermoore
HuggingFace Datasets
Community datasets available for immediate download on HuggingFace.
- scbe-aethermoore-training-data
- scbe-red-team-benchmarks
- polly-training-data
- polly-chat-seed
Prompt Injection → Bit Signatures
24,254 labeled prompts from 4 public injection datasets, each mapped through the Six Sacred Tongues bijective tokenizer into a lossless bit signature. Published as a HuggingFace dataset with stratified train/val/test splits.
- Sources: neuralchemy (6,274), SPML (16,012), jackhhao (1,306), deepset (662)
- Stratified 70/15/15 splits by source × label
- SHA-256 + bit histogram + Shannon entropy + phi-weighted sums
- Baseline: 0.9201 AUC on 31 features (no neural net)
- Apache-2.0 license, immediate download, 44 MB
- Full pipeline + classifier code in scbe-experiments
HuggingFace Model Zoo.
Experimental checkpoints published from SCBE data work. Model-card descriptions identify their intended lane; they are not claims of frontier capability.
polly-base-v1
Early Polly checkpoint fine-tuned on SCBE governance data.
Polly Familypolly-base-v2
Second Polly checkpoint intended to test governance alignment and session handling.
Polly Familypolly-base-v3
Later Polly checkpoint incorporating theory-document and conlang training data.
Polly Familypolly-chat-v1
Chat-optimized Polly. Conversational fine-tuning for interactive sessions.
Polly Familypolly-chat-v2
Improved chat model with better context retention and governance routing.
Polly Familypolly-chat-v3
Latest chat variant. Session transcripts and red team hardening included.
Polly Familypolly-instruct-v1
Instruction-following Polly. Trained on system-level SFT pairs for precise task execution.
Polly Familypolly-instruct-v2
Enhanced instruction model. Governance-aware task routing and multi-step reasoning.
Polly Familypolly-instruct-v3
Latest instruct variant with full 14-layer pipeline awareness.
Polly FamilyPHDM-21D Embedding
21-dimensional Poincare Ball embedding. 6D hyperbolic + 6D phase + 3D flux + 6D audit.
GeometricFoundationGeoSeed Tokenizer
Bijective tokenizer trained on Sacred Tongues vocabulary. Context-aware encoding.
TokenizationSCBE Semantic Projector
Semantic projection model. Maps natural language to 14-layer pipeline coordinates. F1: 0.813.
ProjectorNeed something specific?
Custom dataset work can be scoped around a model architecture, domain, format, and measurable validation plan.