On 14 March 2023, OpenAI introduced GPT-4. The important change was not that every problem had suddenly been solved. It was that the range of tasks for which a general-purpose language model looked professionally useful became materially wider.
GPT-4 made the transition from experimental chatbot to potential work infrastructure easier to imagine.
It did not make human verification optional.
Key takeaways
- GPT-4 widened the professional-use threshold. More complex writing, analysis, coding and document tasks became plausible candidates for AI assistance.
- Benchmark performance was evidence of capability, not proof of professional reliability. Real work still required context, review and accountability.
- The strategic shift was from isolated prompts to workflows. Once a model could contribute across multiple stages of a task, firms could begin redesigning processes around it.
What happened then
OpenAI introduced GPT-4 as a more capable successor to earlier models and documented stronger performance across a range of academic and professional benchmarks.
It also introduced multimodal input capabilities, allowing the model to process images as well as text in supported settings.
Those facts were significant. They suggested that the model was not useful only for generating prose. It could potentially participate in a broader set of analytical and technical tasks.
OpenAI nevertheless continued to document important limitations, including factual errors and reasoning failures.
Interpretation: the threshold moved
ChatGPT had demonstrated that a conversational model could attract mass use.
GPT-4 made a different proposition more credible: that the same broad model family could become a reusable layer inside professional workflows.
That distinction matters.
A novelty is something people try. Infrastructure is something around which processes begin to change.
How the mechanism works
Improved model capability expands the number of tasks that can be delegated partially to software.
A professional can use the model to create a first structure, compare alternatives, transform unstructured information, generate code, interrogate a document or challenge an initial draft.
Each useful stage reduces the marginal cost of another iteration.
The economic effect therefore does not require the model to replace the worker. It can arise because one worker can complete more cycles of research, drafting and revision in the same period.
The constraint moves toward verification.
The strongest countercase
The best objection is that model benchmarks do not measure the whole professional task.
A tax analysis is not merely an answer to a question. A legal conclusion depends on current law, jurisdiction, facts, evidence, interpretation and responsibility. A corporate decision may require access to confidential information and accountability for the consequences.
A model can perform impressively on a benchmark while remaining unsuitable for unsupervised professional judgment.
GPT-4 therefore raised the usefulness threshold without eliminating the reliability problem.
What changed since then?
Later models and tools widened the range of possible workflows again. AI moved increasingly into software development, document processing, research and internal business operations.
But the central distinction remains useful: capability is not accountability.
The more useful the model becomes, the more important it becomes to decide which outputs may be accepted automatically, which require review and which decisions should remain human.
Scenarios, not forecasts
In an augmentation scenario, professionals remain responsible while AI compresses research, drafting and administrative work.
In a workflow-automation scenario, increasingly reliable systems execute multi-step tasks with humans supervising exceptions and approvals.
In a high-stakes ceiling scenario, regulated and consequential decisions continue to require intensive human review even as lower-risk work becomes heavily automated.
Observable error rates, auditability, confidentiality controls and professional rules will determine the boundary.
Practical consequences
The practical question for a firm is not whether a model is “intelligent”.
It is whether a specific workflow can be decomposed into stages with different risk levels.
Low-risk transformation can be automated aggressively. Research may be accelerated but sourced. Drafts can be generated but reviewed. High-impact conclusions require accountable judgment.
For internationally active firms, greater automation can make smaller teams viable across multiple markets. It does not remove local rules governing employment, data, corporate management, tax nexus or regulated activity.
Technology changes the operating model.
It does not suspend jurisdiction.
Sources
- OpenAI, GPT-4 Research, 14 March 2023: https://openai.com/index/gpt-4-research/
- OpenAI, Introducing ChatGPT, 30 November 2022: https://openai.com/index/chatgpt/
- OECD, Generative AI and the SME Workforce, 2025: https://www.oecd.org/en/publications/generative-ai-and-the-sme-workforce_2d08b99d-en.html
Disclaimer
This Insight is general information and analysis, not legal, tax, employment, technology-risk or regulatory advice. AI capabilities, controls and applicable rules evolve quickly and should be assessed against the specific workflow and jurisdiction.
