- How are scenarios selected?
- We start from your policies, architecture, known incidents, and representative requests—then add edge cases and adversarial variants mapped to your risk model.
- Do you need source-code access?
- Usually no. We need enough access to exercise the workflow—staging endpoint, callable function, or representative traces—plus policy and behavior documentation.
- Can you test private or internal systems?
- Yes, when we can reach a controlled test environment and obtain intake materials to define realistic scenarios.
- Can the work be performed entirely in staging?
- Yes. Most baselines run against staging or sandbox environments. Production access is optional and scoped by agreement.
- How much client engineering time is required?
- A technical contact joins scope and intake, helps establish test access, answers focused questions, and attends the findings readout. We design the scenarios, run the evaluation, adjudicate results, and prepare the deliverables.
- What happens if no serious failures are found?
- You still receive the tested coverage, evidence, limitations, and reusable regression suite. A well-supported clean result is useful evidence—not a promise that the system can never fail.
- Are client prompts sent to external model providers?
- Only as required to execute your workflow in the agreed environment. We document subprocessors and can align with your provider data-retention settings.
- Who owns the evaluation dataset?
- You do. The regression suite and scenario definitions are delivered for your ongoing use unless otherwise agreed in writing.
- How do you determine severity?
- By business impact, exploitability, frequency likelihood, and whether the failure violates stated policy or creates unauthorized commitments.
- What happens when evaluators disagree?
- Uncertain outcomes are flagged for human review. Split judgments are documented with rationale—not silently averaged away.
- Does AugmentWorks implement the recommended fixes?
- The baseline delivers evidence and remediation guidance. Implementation is not included. A separately scoped remediation validation can rerun the same scenarios after your team makes changes.
- What does the engagement cost?
- The AI Reliability Baseline is a fixed founding-client fee of $7,500 for one workflow. Follow-on work is quoted after scope confirmation.
- What is explicitly out of scope?
- Infrastructure pentesting, formal certification, legal opinions, unbounded enterprise review, implementation work, and fully automated production monitoring.
- Is this a certification or penetration test?
- No. We deliver an independent assessment with evidence and recommendations—not a badge, compliance certification, or traditional infrastructure pentest.
- What happens after the assessment?
- You receive reports, a regression suite, and an optional separately scoped retest after remediation.