AI for Data Analysis: What the Tools Do, Where They Fail, and the Operator's 7-Step Protocol

AI for data analysis is only as good as the code it runs and the checks you make it pass.
Every operator sits on data they never read. Lead sources in the CRM, check-in scores from clients, refund requests, ad spend exports, a training log going back three years. The spreadsheet exists. The answer is in it. Nobody has the hours to pull it out.
AI closes that gap faster than anything before it. It also produces a confident wrong number faster than anything before it. The difference between the two outcomes is not the model you pick. It is the protocol you run.
This guide covers what the tools actually do as of October 2026, where the published evidence says they fail, and a seven-step protocol, with prompts, that turns a raw export into a decision you can defend.
What AI for data analysis actually means
There are two very different things hiding under the same phrase. The first is a language model reading your numbers and predicting what an answer should look like. The second is a language model writing code, running that code against your file, and reporting what the code returned.
Only the second one is analysis. When Anthropic launched its analysis tool in October 2024, the pitch was exactly this distinction: with code running inside the chat, you get "answers that are not just well-reasoned, but are mathematically precise and reproducible" (Anthropic, analysis tool announcement). The same page now carries an update dated November 5, 2025: "The analysis tool is being replaced by more powerful code execution capabilities." The principle survived the product change. Arithmetic belongs to code, not to prediction.
OpenAI describes the same mechanism on its side. For some data tasks, ChatGPT "writes and runs Python code in a stateful Jupyter notebook environment" (OpenAI Help Center, Data analysis with ChatGPT). Read that sentence carefully. "Some" tasks. Not all of them. Your first job is to make sure the tool is in the code path, not the guessing path.
This matters for a reason older than any chatbot. The NIST engineering statistics handbook describes exploratory data analysis as an approach that postpones assumptions about the model in favour of "allowing the data itself to reveal its underlying structure and model" (NIST/SEMATECH e-Handbook, What is EDA?). A model that predicts plausible numbers does the opposite. It imposes the shape it expects. Code lets the data speak.
What the tools can do as of October 2026
Both major assistants now run code against uploaded files. The details below come from the vendor pages fetched in October 2026. Features change often; check the page before you plan around one.
| Capability | Claude | ChatGPT |
|---|---|---|
| Runs code on your data | Yes. The free plan lists the ability to "Search the web, create files, and run code" | Yes. Python in a notebook environment for some data tasks |
| Spreadsheet output | Creates Excel files with working formulas and multiple sheets | Tables, static charts, some interactive charts |
| Known limits stated by vendor | File creation uses internet access that "may put your data at risk" | Python environment cannot make external web requests or API calls |
| Entry price | Free; Pro $17 per month billed annually or $20 monthly; Max from $100 per month | Varies by plan; not covered here |
Claude's pricing figures come from claude.com/pricing as of October 2026. The file creation details come from Anthropic's announcement of file creation, which says Claude can take raw data and return "cleaned data, statistical analysis, charts, and written insights explaining what matters."
The same Anthropic page gives the most honest line in the whole category: "This feature gives Claude internet access to create and analyze files, which may put your data at risk." If your export contains client health data, payment details or anything under a confidentiality agreement, strip it before it leaves your machine. A column you do not need is a liability you do not need.
Where AI data analysis fails
The failure modes are documented. Learn them before you trust a single chart.
Hard, realistic tasks still beat the agents
DSBench, a benchmark built from 466 data analysis tasks and 74 data modelling tasks drawn from real competitions, found that state-of-the-art systems "struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks" (DSBench, arXiv 2409.07703). That paper was published in 2024 and models have improved since. How much they have improved on that exact benchmark is not established here. The lesson holds anyway: multi-table, long-context, real-world analysis is where errors concentrate, and real business data is all three.
Images of numbers are not numbers
OpenAI states that ChatGPT "may not reliably extract exact values from image-based tables, scanned files, or files with complex visual layouts." A screenshot of a dashboard is not a dataset. A scanned invoice is not a ledger. When exact values matter, export the CSV.
The method can be wrong while the code is right
Code that runs without errors can still answer the wrong question. OpenAI's own help page warns that "ChatGPT may choose an analysis method or chart type that does not match your intent on the first try." An average where you needed a median. A sum over a period where a promo distorted half the window. A sample pulled from the oldest records only. Each of these runs cleanly and lies quietly, which is why ordering bias in sampled data deserves its own check.
Silent defaults in the data
The most expensive errors are not in the model. They are in the data before the model sees it. A missing value filled with a default. A currency column treated as dollars. The model will analyse the broken column with the same confidence it analyses the clean one. The case of a missing exchange rate defaulted to 1 shows how far one silent default travels.
The operator's seven-step data analysis protocol
Run this on every analysis that feeds a decision involving money, hiring or your calendar. It takes 30 to 60 minutes for a single export. The time is the cheapest insurance you will buy all quarter.
- Write the decision first. One sentence: if the answer is X, you will do Y. No decision, no analysis. This stops you from fishing through data until something looks interesting.
- Export clean and strip. CSV or XLSX, never a screenshot. OpenAI's guidance is a fair standard for any tool: "Descriptive column headers in the first row", one row per record, no empty rows splitting the data. Delete every column the decision does not need, especially personal data.
- Force the code path. Tell the model to compute with code and show the code. Prompt: Load this file with code. Before any analysis, report the row count, column names, data types, the date range, and the count of missing values per column. Show the code you ran.
- Reconcile against a number you already know. Pick one total you can verify from another system: last month's revenue from your payment processor, client count from your CRM. If the model's total differs, stop. Find out why before anything else.
- Ask the question, then ask for the method. Prompt: Answer this question: [question]. State the method you used, why you chose it over the alternatives, and every assumption you made about the data. If the data cannot answer the question, say so.
- Attack the answer. Run the same question in a fresh chat and compare. Then prompt: List three ways this conclusion could be wrong given this data. For each, run a check with code and report the result.
- Write the decision log. Three lines in your notes: the question, the answer with the reconciled total, the action taken. Next quarter you rerun the same analysis and compare. That is how analysis compounds.
Steps 5 and 6 are not invented caution. They follow vendor guidance directly. Anthropic's documentation on reducing hallucinations tells you to "Explicitly give Claude permission to admit uncertainty", and describes best-of-N verification, where running the same prompt more than once matters because "Inconsistencies across outputs could indicate hallucinations" (Claude docs, Reduce hallucinations). OpenAI says the same thing in plainer words: "review the generated code, outputs, and assumptions before relying on the result."
Five analyses worth running this month
Start small. Anthropic's own advice on file work is to "Start with straightforward tasks like data cleaning or simple reports" before building toward complex models. These five fit that rule and pay back inside a quarter.
1. Lead source to paid client
Export every lead from the last 12 months with source, date and whether they became a paying client. Ask for conversion rate by source and the median days from lead to payment. Most operators find one channel doing most of the work and one channel consuming most of the attention.
2. Churn by month of tenure
Export client start and end dates. Ask for the percentage of clients lost in month one, two, three and beyond. If losses cluster in a single month, that month is your onboarding problem, and a fix there is worth more than any new lead source.
3. Refund and complaint text
Paste refund reasons or support messages, stripped of names. Ask the model to group them into categories, count each category, and quote two examples per category word for word. The word-for-word requirement keeps the categories tied to what customers actually wrote.
4. Ad spend against real revenue
Join your ad platform export with your payment export by week. Ask for cost per paying customer, not cost per lead. Before you trust any short-window figure, read what "current" really means on a dashboard.
5. Your own training log
This is where the body becomes a business asset you can measure. Export sleep, training sessions and resting heart rate from your wearable alongside your weekly output: calls taken, pages shipped, deep work hours. Ask for the relationship between them, and treat the result as a hypothesis about you, not a law. One person's eight weeks of data proves nothing in general. It can still show you which inputs move your week.
The Stoic rule for numbers
The Stoics separated what is in your control from what is not. Apply that line to data. The model's accuracy is not fully in your control. The checks you run are. The decision you write before you look is. The column you strip before you upload is.
There is a second Stoic habit that fits analysis exactly: rehearse the failure before it happens. Before you act on a number, name the three ways it could be wrong. That is premeditatio malorum applied to a spreadsheet, and the discipline of verifying delegated AI work is the same practice in a different room.
AI is leverage for the operator, not a replacement for judgement. The model reads the file in seconds. You still own what happens next.
Frequently asked questions
Can AI really do data analysis accurately?
Accuracy depends on whether the tool runs code. When Claude or ChatGPT executes code against your uploaded file, the arithmetic is computed rather than predicted, which makes results precise and reproducible. Errors then move to method and data quality: the wrong statistic, a broken column, a biased sample. Always reconcile one total against a system you trust before acting on any result.
Which AI is best for data analysis?
Both Claude and ChatGPT run code on uploaded spreadsheets as of October 2026, so the choice matters less than the protocol. Pick the one you already pay for. Claude's free plan lists running code and creating files; Pro is $17 per month billed annually or $20 monthly. Run the same question in both if the decision is large, and compare.
Is it safe to upload business data to an AI tool?
Treat every upload as data leaving your control. Anthropic states that file creation gives Claude internet access, which may put your data at risk, and advises monitoring chats closely. Remove names, emails, payment details and health data before uploading. Keep only the columns the question needs. For regulated data, check your contracts and your plan's data terms first.
Can AI read data from a screenshot or scanned PDF?
Not reliably. OpenAI states that ChatGPT may not reliably extract exact values from image-based tables, scanned files or complex visual layouts. Any tool reading pixels instead of cells carries the same risk. When exact values matter, export a CSV or XLSX from the source system. A screenshot is fine for a quick read, never for a number you will act on.
How do you stop AI from making up numbers in an analysis?
Force it to run code and show that code. Give it explicit permission to say the data cannot answer the question. Ask it to list how its conclusion could be wrong and test each point. Run the same question in a fresh chat and compare outputs; inconsistency signals a problem. Finally, reconcile one total against a source you already trust.
Build the habit, not the dashboard
The operators who get value from AI for data analysis are not the ones with the most tools. They are the ones who ask one sharp question a week, run it through the same seven steps, and keep a log. Fifty-two verified answers a year compound into a business you actually understand.
The same structure works for the body and the mind: a fixed protocol, a daily check, a written review. If you want to install that operating system in five days, start the free 5-Day Stoic Operator Challenge. Train, think, build, and measure what you build.


