Skip to main content

thangldw thangldw / evaluate-ai-release

Evaluate a RAG system or AI-agent release against an accepted baseline, diagnose regressions, and produce explainable PASS, WARN, or BLOCK evidence. Use for candidate traces, RAG evaluation scenarios, model or prompt changes, release gates, citation regressions, abstention checks, and pre-release...

100

thangldw thangldw / ragops-feature

Implement or review a RAGOps capability or repository contract while preserving evaluation semantics, compatibility, offline behavior, and release evidence.

100

thangldw thangldw / ragops-presentation

Create or review RAGOps website, README, demo, portfolio, or presentation materials using measured evidence and the product visual system.

100

thangldw thangldw / ragops-release

Assess RAGOps release readiness, run its quality gates, and produce owner acceptance evidence without publishing or tagging a release.

100

thangldw thangldw / audit-repository-workflows

Audit a local repository for governance gaps, GitHub Actions trust-boundary risks, and unsafe moderation automation. Use for repository security reviews, pull-request workflow audits, pull_request_target or workflow_run analysis, token-permission checks, policy coverage, and requests for a review...

100

thangldw thangldw / manage-evidence-decisions

Create, review, explain, export, verify, or check the freshness of evidence-backed engineering decisions with Proofline. Use for ADRs, architecture decisions, exact source citations, provenance review, stale-decision checks, portable Decision Evidence Packages, and requests to explain which immut...

75