A new AI tool looks great in a 90-second demo. The demo doesn't show you the data policy, the failure modes, the price after the trial, or how hard it is to leave. This is a due-diligence checklist for a freelancer or small business deciding whether to trust a tool with real work and real client data — a 10-point check, red and green flags, a scoring worksheet, and a worked example on a made-up tool.
A polished demo is not enough when an AI tool will touch client data or business workflows. Check ownership, retention, training use, security documentation, permissions, export options, and the consequences of the tool failing before you connect it.
The 10 things to check
| # | Check | Where to look / what "good" looks like |
|---|---|---|
| 1 | Company identity | A real company name, address, and team are findable. An "About" with no names and no legal entity is a caution. |
| 2 | Data handling | Privacy policy states what's collected, where it's stored, and which sub-processors/model providers see it. |
| 3 | Training on your inputs | Do they train models on what you upload? Is there an opt-out, and is it on by default? |
| 4 | Content ownership & licence | Terms should say you own your inputs and outputs, and grant the vendor only the licence needed to run the service. |
| 5 | Retention & deletion | Can you delete data and your account? Does deletion actually remove it, and within a stated window? |
| 6 | Integrations & permissions | When it connects to email/drive/calendar, does it ask for read-only where possible, or full access to everything? |
| 7 | Security posture | Look for specifics (encryption in transit/at rest, SSO, an audited compliance report you can request) — not just the word "secure". |
| 8 | Pricing & trial terms | Price after the trial, whether a card is required up front, auto-renew terms, and how cancellation works. |
| 9 | Export & lock-in | Can you get your data and work out in a standard format if you leave? Or is it trapped in their UI? |
| 10 | Model/provider dependency | Which underlying model powers it? If that provider changes pricing or access, does the tool break? |
Then test the behaviour, not just the paperwork
- Low-stakes first. Run it on a draft only you'll see, or a task with no deadline, before anything client-facing.
- Provoke a failure. Feed it something ambiguous or outside its scope. Does it flag uncertainty, or state a wrong answer in the same confident tone as a right one? Quiet, confident failure is the dangerous kind.
- Check the workflow fit. Does it remove a bottleneck you actually have, or is it an impressive feature attached to no real problem? The second kind gets abandoned in a month.
Red flags
- No company name, no legal entity, no way to contact a human.
- Privacy policy is generic boilerplate that never names a data location or sub-processor.
- Training on your data is on by default with a buried or missing opt-out.
- Terms claim a broad licence to your content beyond running the service.
- No export. Your work only lives inside their app.
- Card required for a "free" trial with auto-charge and an awkward cancellation path.
- Security page is adjectives ("bank-grade", "military-grade") with zero specifics.
- Requests full account access when read-only would do.
- Roadmap and pricing change frequently with no changelog.
Green flags
- Named company, clear docs, responsive support channel.
- Privacy policy names storage regions, sub-processors, and the model provider.
- Training opt-out that's either off by default or a single clear toggle.
- Plain statement that you keep ownership of inputs and outputs.
- One-click export in a standard format; documented account deletion.
- Granular, least-privilege permissions on integrations.
- Transparent pricing page, trial without a card, self-serve cancellation.
- A public changelog and status page.
Scoring worksheet
Tool: __________________________ Date checked: __________ Score each 0 (fail) / 1 (partial) / 2 (good): [ ] 1. Company identity clear [ ] 2. Data handling documented [ ] 3. Training opt-out (off by default = 2) [ ] 4. You own inputs/outputs [ ] 5. Retention & deletion stated [ ] 6. Least-privilege integrations [ ] 7. Concrete security details [ ] 8. Pricing & trial terms clear, no card trap [ ] 9. Export exists, low lock-in [ ] 10. Model dependency understood Behaviour: [ ] Tested on low-stakes work [ ] Failure mode is visible, not silent [ ] Solves a real bottleneck I have Total: ___ / 26 Data sensitivity of intended use (low / medium / high): ______
When to skip the tool
- Any score of 0 on checks 2, 3, 4, or 5 and you'd feed it client data.
- It fails silently and you can't build a reliable review step around it.
- No export and it would hold work you can't afford to lose.
- It duplicates a tool you already have (see avoiding AI tool overload).
- The problem it solves would be fixed by a free setting change elsewhere.
Worked example — a hypothetical tool
Hypothetical, invented for illustration. "InboxZero AI" promises to draft all your client email replies. On the checklist: company is a named entity with docs (2). Privacy policy names storage region but not the model provider (1). Training on inputs is on by default, opt-out is in an account sub-menu (0). Terms confirm you own outputs (2). Deletion is documented, 30-day window (2). It requests full Gmail access, no read-only option (0). Security page lists encryption and offers a compliance report on request (2). Free trial needs a card, auto-renews monthly, cancellation is self-serve (1). Export is copy-paste only (0). Built on a single third-party model with no stated fallback (1). Behaviour: tested on a personal draft, it invented a meeting time that was never discussed and stated it plainly — a silent-ish failure (0).
Total ≈ 11/26, and it would touch a client's entire inbox. Verdict: don't connect it to real accounts. It might be usable as a standalone draft assistant you paste into — never wired into live email — and only after turning off input training.
After it passes
Approval isn't permanent. Models update and behaviour drifts, often with no announcement. Keep a one-line log per tool (name, date checked, score, notes) and re-run the check if output quality shifts or the pricing/terms change. This pairs with our rundown of common AI automation mistakes that cost freelancers clients.