Agent reliability artifacts
Lint manifests, replay tool calls, normalize traces, and export portable run cards.
AuraOne Open
Loading products, source links, and current releases.
AuraOne Open / Reliability evidence
Source-inspectable CLIs, SDKs, and Actions that produce portable evidence. Agent traces. Dataset cards. Robotics quality records.
For engineering, evaluation, and audit teams that need evidence to stay portable across CI systems and providers.
What you can do
Get started
Browse the current public repositories, then select the package or Action that matches the evidence artifact your workflow needs.
gh repo list auraoneai --limit 100Lint manifests, replay tool calls, normalize traces, and export portable run cards.
Version rubrics, protect regression banks, detect contamination, and publish decision-ready results.
Write dataset cards with rights references. Run quality checks and produce a reviewed release manifest.
See how the product moves from setup through a useful output.
01
Choose the exact record your workflow requires. Nothing more.
Artifact definition
02
Read the source, the license, and the permissions it asks for. Check what it sends over the network.
Source and dependency review
03
Execute against synthetic or approved data in the intended repository or CI environment.
Observed runtime record
04
Check the schema, the provenance, and the checksum before you trust it.
Validation report
05
Retain the exact artifact version and source beside the approval or release gate.
Artifact version and source
Overview
Review the current source, release, runtime, and support information.
Availability
Review the current version, source, packages, downloads, supported platforms, and install options.
The Trust Toolkit catalog is current across its public repositories and source-only projects. That covers its PyPI packages, npm packages, and GitHub Actions. Nothing is shared between them. Each tool keeps its own version and install path. Each keeps its own license and runtime boundary.
Toolkit records
Every record shows its current source, public version, and install path. The license and the runtime boundary are stated.
27 open-source products
Validate rubrics, score approved fixtures, audit reviewer agreement and leakage, and generate portable evaluation reports.
One local Python CLI keeps rubric checks, deterministic scoring, reviewer QA, leakage checks, and evidence reports in the same inspectable workflow.
PyPI 0.3.0 is available with a linked source release.
python -m pip install auraone-evalkitRun AuraOne evaluations for an exact commit and publish one lifecycle-aware Check Run plus one idempotent evidence summary.
Check state, evaluated commit, threshold, template evidence, remediation, permissions, and deployment identity remain attached to one native review path.
npm 0.2.0 is available with a linked source release.
npm install @auraone/github-appAgent engineers automating trace ingest, replay, comparison, and CI export.
MIT source
Run Agent Studio's trace and regression workflows in terminals and CI without the visual workbench.
CLI output and exported evidence follow the same trace, replay, and release contracts as the desktop application.
PyPI 0.2.1 is available with a linked source release.
python -m pip install auraone-agent-studio-openReview event streams, failures, and decisions in one self-contained offline HTML artifact.
The canonical viewer exports review metadata without exporting observations, actions, sensor payloads, media, or training shards.
Offline artifact release 0.3.0 is available.
python -m http.server --directory robotics-reviewkit/viewerBrowse, filter, validate, and reproduce public synthetic failure cases.
Each case keeps the symptom, reproduction, expected behavior, observed behavior, tags, and contribution provenance together.
PyPI 0.2.1 is available with a linked source release.
python -m pip install failure-galleryScore repository responses against an AuraOne rubric and publish one consistent GitHub evidence surface.
Annotations, job summary, outputs, threshold decision, and optional idempotent pull-request comment agree on one result.
GitHub Action release 0.2.0 is available.
uses: auraoneai/evalkit-action@v0.2.0Validate required Datasheet, Model Card, and Data Card sections and surface actionable review findings.
Blocking omissions and high-recall PII-like review signals are separated into one inspectable decision and one idempotent summary.
PyPI 0.2.1 is available with a linked source release.
uses: auraoneai/datasheet-ci@v0.2.1Diff and lint rubric changes, annotate affected lines, and publish one merge decision for the exact commit.
The bot updates one Check Run and one marker-owned summary instead of producing duplicate or detached review noise.
Version 0.2.0 is available from the project source.
git clone https://github.com/auraoneai/rubric-pr-bot.gitValidate, lint, diff, and convert portable AuraOne Rubric Schema v1 files.
One criterion-level schema preserves scoring intent across EvalKit and supported framework adapters.
PyPI 0.1.2 is available with a linked source release.
python -m pip install rubric-specMeasure inter-annotator agreement with modern metrics, uncertainty, ordinal support, and missing-data handling.
Agreement estimates and bootstrap intervals are available from one pure-NumPy library with deterministic fixtures.
PyPI 0.1.2 is available with a linked source release.
python -m pip install iaa-kitProbe judge behavior for position, verbosity, self-preference, paraphrase, anchoring, and calibration failures.
Synthetic diagnostic cases isolate judge failure modes before they are mixed into production evaluation scores.
PyPI 0.1.2 is available with a linked source release.
python -m pip install judge-benchValidate and render a disclosure record for judge prompts, calibration, bias, limitations, and intended use.
The card keeps operational judge evidence beside the evaluation system instead of burying it in prose.
PyPI 0.1.2 is available with a linked source release.
python -m pip install judge-cardTranslate one rubric-spec rubric and run configuration into supported evaluation-framework inputs and exports.
The portable rubric remains authoritative while framework-specific files are generated and reviewable.
PyPI 0.1.2 is available with a linked source release.
python -m pip install eval-adapterRun executable rubric-spec v1 compatibility checks and emit reproducible conformance evidence.
Compatibility claims are tied to executable fixtures and a versioned badge rather than documentation alone.
PyPI 0.1.2 is available with a linked source release.
python -m pip install eval-conformance-suiteCreate and validate a portable envelope for evaluation provenance, environment, artifacts, result digests, and optional signatures.
The decision record links code, data, rubric, judge, contamination, and result evidence through one schema.
PyPI 0.1.2 is available with a linked source release.
python -m pip install eval-run-manifestScan evaluation data for n-gram overlap, canaries, answer patterns, hashes, and optional embedding similarity.
Multiple local contamination signals are reported together with evidence instead of collapsed into one unsupported certainty score.
PyPI 0.1.2 is available with a linked source release.
python -m pip install contamination-auditCompare before and after prompt and rubric directories and generate deterministic drift review notes.
Scoring-boundary, criterion, weight, and prompt changes are surfaced as evidence-oriented review items.
PyPI 0.1.7 is available with a linked source release.
uses: auraoneai/prompt-rubric-drift@v0.1.7Reviewer-quality teams testing agreement and adjudication pipelines.
MIT source
Generate controlled synthetic annotator disagreement before real reviewer data is collected.
Deterministic disagreement patterns let teams test metric, missingness, severity, and adjudication behavior against known conditions.
PyPI 0.1.2 is available with a linked source release.
python -m pip install synthetic-disagreementTurn one agent trace into portable Markdown, HTML, and JSON evidence cards.
Goal, outcome, tools, retries, data touched, checks, timing, and remediation remain attached to the same run identity.
PyPI 0.1.2 is available with a linked source release.
python -m pip install agent-trace-cardNormalize agent tool-call traces into deterministic replay artifacts and pytest cases.
Recorded tool inputs and outputs can be locked, redacted, replayed, diffed, and promoted into CI without rerunning the live tool.
PyPI 0.1.1 is available with a linked source release.
python -m pip install tool-call-replayConvert OpenTelemetry or Phoenix GenAI trace exports into local evaluation regression cases and manifests.
Operational traces become portable evaluation inputs without requiring a proprietary trace store.
PyPI 0.1.2 is available with a linked source release.
python -m pip install otel-eval-bridgeScan MCP manifests, package metadata, source, and docs for tool, auth, filesystem, shell, network, and readiness risks.
The linter reports inspectable repository evidence and configurable severity instead of claiming runtime security certification.
PyPI 0.1.6 is available with a linked source release.
uses: auraoneai/mcp-risk-linter@v0.1.6Validate agent cards, capabilities, structured payloads, errors, cancellation, and task lifecycle behavior.
Offline contract fixtures test the public protocol boundary without requiring a live agent deployment.
PyPI 0.1.5 is available with a linked source release.
uses: auraoneai/a2a-contract-test@v0.1.5Validate and render structured cards for robot morphology, sensors, control, environment, data, and known limitations.
Embodiment-specific release facts stay machine-readable and render into a reviewer-facing card.
PyPI 0.1.2 is available with a linked source release.
python -m pip install embodiment-cardValidate and summarize human intervention and recovery segments in robot episodes.
Recovery latency, success, recurrence, operator action, and evidence references stay attached to the source segment.
PyPI 0.1.2 is available with a linked source release.
python -m pip install robot-recovery-benchGenerate simulator-light perturbation diagnostics for vision-language-action policies.
Task-instruction and observation perturbations produce bounded, reviewable diagnostics without claiming full simulator coverage.
PyPI 0.1.2 is available with a linked source release.
python -m pip install vla-robustness-kitRobotics platform engineers needing Robotics Studio workflows without the desktop UI.
MIT source
Index robotics datasets, generate thumbnails, run QA and clustering, and export reviewed evidence through a headless engine.
The CLI exposes the same dataset and evidence contracts used by Robotics Studio Open for automation and integration.
PyPI 0.1.2 is available with a linked source release.
python -m pip install robostudio-engineContinue exploring
Each one lists its source, current version, and install path.