Participants: start with START_HERE.md.
A four-day, instructor-led, hands-on enterprise AI programme built around one connected synthetic scenario: Asset A-001, a cooling-water pump.
The programme uses Databricks as the primary hands-on environment and progressively connects enterprise data, machine learning, Generative AI, Retrieval-Augmented Generation (RAG), AI agents, governed workflows and a final integrated application.
Training notice: every asset, document, procedure, observation and dataset in this repository is synthetic and designed only for training. Nothing here is operational engineering guidance.
- Create a free Databricks account: https://login.databricks.com/signup
- Open the training Release: https://github.com/Decoding-Data-Science/enec2026/releases/tag/enec-2026-final
- Download
ENEC_2026_V2_2_Data_and_PDF_Corpus.zip. - Extract the ZIP on your computer.
- Confirm you can see
nuclear_enterprise_360_v2_2_clean.dband the nine A-001 PDFs. - In Databricks, open/import this repository and run:
Day_1_Databricks_Data_ML/notebooks/00_Setup_Nuclear_Enterprise_360.py - Run the first setup cell. It creates the schema and
training_filesVolume. - In Databricks choose New → Add or upload data → Upload files to a volume.
- Upload
nuclear_enterprise_360_v2_2_clean.dbto:/Volumes/workspace/nuclear_enterprise_360/training_files/ - Return to the setup notebook and run the upload-check cell, then run the remaining setup cells.
- Verify A-001:
SELECT *
FROM workspace.nuclear_enterprise_360.asset_360
WHERE asset_id = 'A-001';- Verify the hourly teaching dataset:
SELECT COUNT(*)
FROM workspace.nuclear_enterprise_360.a001_sensor_hourly;Expected result: 8,760 rows.
The
.dbfile is the actual SQLite training database. Do not try to open it as a text file or notebook. The setup notebook converts it into Databricks Delta tables.
A-001
↓
Databricks & Enterprise Data
↓
Machine Learning / Forecasting
↓
Generative AI
↓
Enterprise RAG
↓
AI Agents
↓
Skills + Memory + Progressive Disclosure
↓
Governed / Multi-Agent Workflows
↓
Integrated Enterprise AI Application
↓
Human Review & Oversight
| Folder | Purpose |
|---|---|
| Day_1_Databricks_Data_ML | Databricks foundations, enterprise data exploration, vibration forecasting and optional anomaly detection |
| Day_2_GenAI_RAG | LLM foundations, document ingestion, chunking, embeddings, retrieval, authority filtering and RAG |
| Day_3_Agentic_AI | Tool use, SQL + RAG agents, skills, progressive disclosure, memory and governed agent patterns |
| Day_4_Integrated_Capstone | End-to-end integration, governed orchestration, capstone, evaluation and participant demos |
| data | V2.2 database package, data dictionary, validation script and quick-start SQL |
| resources | Programme agenda, participant setup, instructor run-of-show, slides/archive catalog and additional references |
| reference | Earlier complete Databricks kit retained for instructor/reference use |
Participants work with A-001 throughout the programme.
Structured sources include:
- asset master and health information
- sensor readings
- inspections and maintenance history
- work orders
- projects and risks
- document metadata and links
Unstructured sources include:
- enterprise reliability policy
- operating guide
- approved, superseded and draft procedures
- troubleshooting guide
- condition-monitoring report
- work order
- field-condition report
The teaching point is not simply to retrieve the most similar or newest information. The system must reason about authority, approval status, evidence quality, provenance and human oversight.
Primary target:
- Catalog:
workspace - Schema:
nuclear_enterprise_360 - Volume:
training_files - Volume path:
/Volumes/workspace/nuclear_enterprise_360/training_files
Recommended workspace folders:
/Workspace/Nuclear_Enterprise_360/
├── A001 Documents/
├── Agent_Skills/
└── Notebooks/
- Open START_HERE.md.
- Clone/download this repository.
- Download the V2.2 data/PDF companion pack from GitHub Releases.
- Run the Day 1 setup notebook in Databricks.
- Follow the README inside each day folder.
The repository now also includes readable mirrors of all nine A-001 documents under Day_2_GenAI_RAG/documents/a001/.
- Read resources/INSTRUCTOR_RUN_OF_SHOW.md.
- Validate V2.2 with
data/verify_v2_2.py. - Keep the pre-staged Databricks environment as the live-demo fallback.
- Use the slide archive in
resources/slides/. - Keep the same A-001 narrative across all four days.
The primary training database is V2.2 Clean Beginner Version, created 16 September 2026.
Core forecasting flow:
- raw:
sensor_readings_a001_1min— 525,600 rows - teaching view:
a001_sensor_hourly— 8,760 hourly rows - recommended target:
avg_vibration_mm_s - recommended forecast horizon: 7 days / 168 hours
See data/README.md for details.
- One connected scenario rather than unrelated daily examples.
- Practical application before deep platform administration or theory.
- Build capability progressively; do not reveal the full architecture too early.
- Keep data, knowledge, tools, skills, memory and authority conceptually separate.
- Treat relevance and authority as different questions: relevant does not automatically mean trusted.
- Preserve uncertainty and evidence gaps.
- Human reviewers retain final operational authority.
By the end of Day 4, participants should be able to explain and demonstrate how a governed enterprise AI application can combine:
- structured enterprise data
- predictive ML outputs
- trusted organisational documents
- RAG with source attribution
- tools and agent routing
- workflow state and memory
- governance, auditability and human review
Decoding Data Science — Learn by Building
ENEC 2026 synthetic enterprise AI training repository.
\n## Binary companion packs
Two large/binary companion packs are distributed through GitHub Releases:
ENEC_2026_V2_2_Data_and_PDF_Corpus.zipENEC_2026_Slides_and_Trainer_Materials.zip
See resources/BINARY_ASSETS.md.
For repository-owner publishing steps, see resources/OWNER_FINAL_STEPS.md.