Source identity
Normalized objects retain checksums, citations, safety state, and corpus fingerprints through retrieval.
Project Brief
The repository normalizes synthetic text, transcript-derived, OCR-derived, and Markdown material into provenance-rich objects, gates retrieval inputs for safety, produces local lexical/vector records, and exposes portable hybrid retrieval logic used by a public cited-answer demo.
Live Demo
Open the standalone surface to see the published corpus status, ask a question, and inspect citations without scrolling through the project narrative.
Implemented Pipeline
The pipeline preserves source identity, checksums, citations, safety state, and corpus fingerprints through retrieval.
Privacy and answer-grounding gates block unreviewed inputs and require cited, bounded responses.
Evidence
Normalized objects retain checksums, citations, safety state, and corpus fingerprints through retrieval.
Transcript and OCR-style examples become cleaned, citation-preserving corpus records after public-safety review.
Generated lexical and vector records support portable hybrid retrieval for a public cited-answer demo.
Privacy and answer-grounding gates block unreviewed inputs and require cited, bounded responses.
Serving Layer
The repository exposes portable hybrid retrieval logic over the same safety-reviewed records used by the public cited-answer demo.
Projects Qualification
Strong general AI systems alternate with unusually clear provenance and governance evidence; less education-specific than the top three.
Safety Boundary
Public examples must use synthetic, sanitized, public-domain, or clearly licensed source material. Raw private artifacts, private source paths, credentials, course identifiers, and copyrighted lecture text stay outside the public repository and public RAG index.