# Test local and cloud AI on your own tasks

Neaptide worksheet · September 27, 2026

This is a proposed method for a small working pilot, not an industry standard or a ready-made benchmark. Enter results after testing. Do not send an option any data it is not permitted to process.

## 1. Conditions and decision

- Decision to make:
- Task and typical monthly workload:
- Current method and its costs:
- Data permitted for each option:
- Unacceptable errors:
- Person responsible for accepting results:
- Minimum acceptable quality and timing (set before testing):
- Comparing models under matching conditions, or complete workflows:
- Pilot date and duration:

## 2. Configurations

| Parameter | Local option | Cloud option |
| --- | --- | --- |
| Model name and version | | |
| Application / API / mode | | |
| Quantisation, if relevant | | |
| Hardware / plan | | |
| Context and generation settings | | |
| Available tools and search | | |
| Where documents are read and indexed | | |
| Retention and training terms | | |
| Network, concurrent users, other applications | | |

## 3. Task card — copy for each example

- ID:
- Source material and permission to process it:
- Exact prompt:
- Correct facts and their source, or a behaviour check:
- Acceptance criteria:
- What the model should do when information is missing:
- Critical errors:

Include routine tasks, a question the materials cannot answer and a case with conflicting data. Keep the answer key separate from model input. You can start with 12 tasks, but the mix should reflect your work.

## 4. Results log — one row per attempt

| ID | Option and run | Time to first visible response, s | Time to complete response, s | Review and corrections, min | Error | Accepted against the criteria? | Usage cost |
| --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | |

Record whether the model was already loaded. Do not combine technical metrics from different applications without checking their definitions. Keep failed attempts in the log. If corrections lead to an accepted result, count all the time spent.

## 5. Costs for the same workload and period

- Additional hardware, less any residual value counted / useful life:
- Initial setup / the same period:
- Service, licence and integration fees:
- Additional energy:
- Maintenance and administration (hours × hourly cost):
- Review and corrections (hours × hourly cost):
- Other expenses:
- Shared costs excluded from both options:
- Total accepted results:

Cost per accepted result = costs for the whole compared workload, including failed attempts / number of accepted results. With zero accepted results, the metric is undefined.

Do not generalise a small pilot to all future tasks. Test long documents, peak demand and model version changes separately.

## 6. Pilot decision

- Chosen option and reasons:
- Tasks the conclusion applies to:
- What did not pass:
- Remaining limitations:
- When to repeat the test:
