Skip to content
AugmentWorks

Primary engagement

AI Reliability Baseline

A fixed-scope assessment of one production AI workflow.

The situation it addresses

Your team may already test prompts manually, review a few traces, and maintain a small set of golden examples. That can verify expected behavior, but it often misses adversarial wording, retrieval errors, policy edge cases, multi-turn pressure, and unsafe tool decisions.

The baseline turns one important workflow into a structured evaluation program—with evidence your engineers and security stakeholders can act on.

Good fit if

  • A staging endpoint, callable function, or representative traces exist
  • The team can define expected and prohibited behavior
  • A technical owner can answer focused intake questions
  • There is an upcoming launch, migration, enterprise review, or known concern

What counts as one workflow

  • Answering refund questions from retrieved policy documents
  • Reviewing a contract and returning structured risk fields
  • Triaging support tickets and updating a CRM
  • Answering employee questions from internal documentation
  • Extracting values from uploaded financial documents

Scope boundary

The fixed fee depends on a narrow, written boundary. Scope is confirmed before Day 1.

  • One defined user journey or decision path
  • One agreed staging environment or trace set
  • One expected-behavior policy set
  • Up to 30–50 targeted scenarios
  • One findings readout and delivered regression suite
  • Additional workflows require separate written scope

What you provide

  • Staging endpoint, callable function, or representative traces
  • Relevant policies and expected-behavior documentation
  • Known incidents or concern areas
  • Representative user requests and edge cases
  • A technical point of contact for intake questions

What you receive

Risk and behavior map

What the system is permitted, required, and prohibited from doing in scope.

Custom evaluation set

30–50 reusable scenarios covering ordinary behavior, edge cases, and adversarial pressure.

Human-adjudicated findings

Severity-ranked failures with transcripts, expected behavior, observed behavior, and rationale.

Engineering remediation plan

Recommended changes to prompts, retrieval, guardrails, tool authorization, or workflow design.

Regression package

Machine-readable tests suitable for future releases and model migrations.

Executive and engineering reports

Separate materials for decision-makers and implementers, plus a findings readout.

Timeline & pricing

15 business days after intake materials are complete: scope confirmation → intake → scenario design → execution → human adjudication → reports → readout.

Founding-client fee: $7,500 fixed for one workflow in scope.

Request a 20-minute fit call

Explicitly out of scope

  • Infrastructure penetration testing
  • Formal certification or compliance opinions
  • Unbounded review of an entire product surface
  • Implementation of recommended changes
  • Retesting after remediation (available under separate scope)
  • Fully automated production monitoring
  • Legal or regulatory advice

What happens next

  1. Submit the contact form with your workflow and primary concern.
  2. We respond within two business days to schedule a 20-minute fit call.
  3. If aligned, we send a statement of work with scope, access, and timeline.
  4. After signature, you receive portal access for intake and deliverables.