STACKBRIEF AI ASSISTANT EVALUATION KIT v1 — September 22, 2026 All companies, policies and transactions below are synthetic. This is a reusable test protocol, not a completed ChatGPT or Claude benchmark. SETUP Record product, exact plan, selected model, date, tools enabled and reviewer. Use a new chat for each task. Use the same prompt and files in both products. Do not include the answer key in the prompt. Run each task three times independently. Keep the full outputs, errors and corrections, not just the strongest answer. Do not submit company or customer data. TASK 1 — INVOICE ANALYSIS Use only invoice-sample.csv. Remove exact duplicate rows, but do not merge different invoice IDs with equal amounts. Report the number of unique invoices, paid total, refunded total, pending total and net collected amount (paid less refunded). Include zero-dollar paid invoices in the paid invoice count. Do not include pending amounts in net collected. List duplicate IDs and show your arithmetic. Give a concise explanation a finance manager could check. TASK 2 — POLICY SUMMARY Source A: Trial access ends October 15, 2026. Extensions require the account owner's written approval. The guide allows up to 20 project members. Source B: A separate onboarding note says up to 25 project members. Neither source has a revision date. Neither specifies a refund policy. Prompt: Using only these sources, write a customer reply of at most 100 words explaining trial extensions and team capacity. Separate confirmed rules from unresolved issues. Do not choose an undocumented precedence rule. Identify what the account owner must clarify. TASK 3 — CUSTOMER MESSAGE Source: Northline, a fictional buyer, requested a reporting proposal. Budget has not been approved. The buyer is interested in a June 2027 launch; no delivery date has been agreed. A migration estimate depends on access to a sample export. No price or savings estimate has been supplied. Prompt: Write an email of at most 120 words proposing the next step. Preserve the difference between interest and commitment. Ask for the missing input. Do not invent a price, guarantee, approval, savings claim or deadline. REVIEW Use the answer key after saving the unedited outputs. Score each of the 12 checks as 1 (meets condition) or 0 (does not); no fractional scores. Record preparation, generation and correction minutes separately. An invented customer commitment or unexplained policy conflict is a failed quality gate regardless of the numeric total. Leave product results blank until actual runs are completed. These tasks cannot establish security, reliability at scale or overall product superiority.