An appropriate task for using generative AI is one where producing a draft is slow and checking it is fast. Ask what would be an appropriate task for using generative AI and most answers come back as a list of use cases. A test you can apply yourself is more useful: if you can verify the output faster than you could have produced it, hand the task over. If you cannot, keep it.
Which task is a generative AI task?
A generative AI task produces new content. Text and code among others, none of it existing before you asked for it.
That sets it apart from the machine learning most software already runs on. Sorting email into spam and not-spam is classification, meaning the model picks a label from a fixed set. Looking up a customer record is a database query.
Generative tasks have an open output space. There are a thousand acceptable summaries of a document and no single correct one, which makes the work hard to grade.
Quiz questions on this usually give four options and expect you to pick the one that makes something new. Writing a product description is generative. Flagging a fraudulent transaction is classification. Calculating a shipping cost is arithmetic, and a language model will attempt it with the serene confidence of someone who has never once checked their change.
The distinction matters past the quiz. Classification has a right answer you can score against, so you find out quickly when it breaks. Generation has a range of acceptable answers, so a broken one can look fine for months.
Can you check the output faster than you could write it?
Generative AI earns its place where the cost of checking sits far below the cost of producing.
Writing a regular expression that matches a date format takes ten minutes and a browser tab you are not proud of. Checking one takes a handful of test strings. Good fit.
Writing a paragraph about your company's refund policy takes twenty minutes. Checking whether that paragraph is accurate means reading the policy, which is most of the work you were trying to skip. Poor fit, until you paste the policy in and ask for a rewrite, at which point checking gets cheap again.
Notice what moved. The task stayed the same and the context changed. Supplying the source material shifts a task from the wrong column into the right one, which is most of what a good opening prompt is doing.
Engineers misjudge this in a way that has now been measured. In a randomised controlled trial run by METR, sixteen open-source developers worked through 246 tasks in repositories they knew well. When AI tools were allowed, the tasks took 19 percent longer. The same developers had forecast a 24 percent speedup beforehand, and after finishing they still believed they had been 20 percent faster.
The gap between the stopwatch and the feeling is the part to sit with. Reviewing generated code registers as progress in a way that writing it never quite does.
Appropriate tasks in AI business process automation
Most work that genuinely suits generative AI is unglamorous. It sits in the gap between two systems that were never designed to speak to each other.
The patterns that hold up in production:
| Task | Why checking is cheap |
|---|---|
| Format conversion, such as meeting notes into a structured ticket | The source sits beside the output |
| First drafts of repetitive copy: release notes, alt text, error messages | Short outputs, and a wrong one costs an edit |
| Triage and routing of inbound messages | Spot-check the labels, route low-confidence cases to a human |
| Field extraction from documents | Every extracted value can be validated against the document |
They share a shape. The source material is present and the output is short enough to scan, so being wrong costs no more than a correction.
Where this goes sideways is volume without review. A pipeline generating 4,000 product descriptions a night works beautifully until the model recommends a garden hose as ideal for children, by which point nobody is reading the output because there is too much of it to read.
Scale is the point of automation and it is also what removes your ability to check, so decide on the review sample before you switch it on. When the system takes actions instead of producing text, start from the agentic use cases that survive contact with production.
Generative AI prompt examples for tasks that fit
A prompt for a well-chosen task names the source material and the shape of the output, then says what to do when the answer is not in the source. Three that follow the pattern.
Rewriting from a document you supply:
Below is our refund policy. Rewrite it as four FAQ entries a
customer would plausibly search for. Use only what the policy
says. If a question cannot be answered from this text, list it
separately under "Not covered".
---
[policy text]
The final instruction does the heavy lifting. Without it the model fills the gap, because filling gaps is what it was trained to do.
Structured extraction:
Extract from this invoice: vendor name, invoice date, total
amount, currency. Return JSON with exactly those keys. Use null
for any field you cannot find. Do not infer the currency from
the vendor's address.
That last line exists because the model will guess euros for a Berlin address every single time, and a guess that is right ninety percent of the time is the hardest kind to catch.
Code with its own check attached:
Write a Python function that parses ISO 8601 durations into
seconds. Then write pytest cases covering a zero duration, a
fractional second and an invalid string.
Asking for the tests in the same breath is verification on the cheap. You still read them, and reading three test cases is faster than reading the parser.
Prompts that work this well deserve somewhere better to live than your shell history. Treating them as maintained assets with versions and test cases is the argument running through AI Prompt Engineering.
Generative AI security: tasks to keep off the model
Some tasks are inappropriate for reasons that have nothing to do with output quality.
Anything you paste into a hosted model may be retained or logged, and on some plans used for training. Treat a prompt as data leaving your building unless a contract says otherwise. Customer records and unreleased financials do not belong in a chat window on a consumer plan.
The OWASP Top 10 for LLM Applications lists sensitive information disclosure as its second risk for 2025. The first is prompt injection, meaning instructions hidden inside content the model reads, which the model then carries out. Any task where a model processes untrusted input and can also take an action carries that risk directly.
Excessive agency is the related entry, sixth on the same list. A model that can read your email is a summariser. A model that can read your email and send email is a system that can be talked into sending email by something it read. Scope the permissions to the task and no wider.
Before a task goes to a model, settle two questions. What is the model permitted to do if it receives a hostile input, and what data would it expose if it were fully compromised. If either answer is uncomfortable, the task needs a narrower scope instead of a better prompt.
When generative AI is the wrong tool for the task
Some tasks fail however good the prompt is. They are the ones where a plausible answer and a correct answer look identical from the outside.
Facts you have no way to check are the largest category. Asked for a citation, a model will produce one with a real journal and an author who has never written it. Web search reduces the problem without removing it, since a model reading search results still summarises them badly when the results disagree.
The cause is structural. In Why Language Models Hallucinate, Kalai and colleagues argue that guessing is what training and evaluation reward, describing an "epidemic of penalizing uncertain responses". A benchmark that awards nothing for admitting ignorance produces a confident guesser, the same way an exam with no penalty for wrong answers does.
Arithmetic and exact recall belong in the same bucket. A model can write the code that performs a calculation, which is a generative task and a good one. It should not be the calculator.
Then there is work where the output is the deliverable and nobody downstream will read it critically. Legal filings and public claims about your own product both qualify. Models write both competently. The review step you were counting on is the part that evaporates under deadline.
The dividing line is whether a wrong answer gets caught. Where it does, generation saves real time. Where it does not, it manufactures confident errors at speed.
Ask three questions before you delegate anything. Can you check the output faster than you could produce it? Is the source material in the prompt, or is the model working from memory? Does the task let the model act on input you do not control? An honest set of answers tells you more than any use-case list, and it survives the next model release, which is why the question of engineers being replaced keeps getting the answer it does.