I use AI in my actuarial practice every week. It runs a monthly cash-flow file process for one of my clients that used to mean hours of pivot tables and copy-paste, and it writes much of the code behind my own tools. I still would not let it sign anything. Both of those things are true at once, and the gap between them is the real question about AI in our work.

The question "will AI replace actuaries" has been asked so many times it no longer produces information. It is a headline, not a research program. The more useful question is the one the profession's own research bodies are now putting money behind: which specific parts of actuarial reasoning a model can be deliberately engineered to support, and where AI still fails badly enough that a human signature has to stay on the work. The Casualty Actuarial Society (CAS) is running exactly that inquiry in public, and its framing tells you more about the real state of play than any think-piece.

What the CAS is Actually Asking

In early 2026 the CAS Artificial Intelligence Working Group issued a Request for Proposals (RFP) on adapting Large Language Models (LLMs) for specialized Property and Casualty (P&C) actuarial reasoning. The wording repays close reading. The RFP explicitly distinguishes prior work, which treated general-purpose models as tools that perform calculations or respond to prompts, from what it now wants to fund: research into the mechanisms by which a model's behavior is shaped for actuarial use. It asks for approaches that move beyond out-of-the-box or prompt-driven applications, and toward models that are structured, trained, or constrained to reflect actuarial logic, data structures, and judgment, delivered as a reproducible system a practitioner could actually adopt. The candidate techniques it names include fine-tuning, structured context engineering, retrieval-augmented or modular architectures, and training domain-specific models from scratch.

That is a precise and revealing ask. The CAS is not commissioning research to find out whether ChatGPT can pass an exam. It has moved past that. The working assumption behind the RFP is that a general model prompted in the ordinary way is not adequate for core actuarial reasoning, and that the interesting research question is what scaffolding, constraint, and domain adaptation it takes to make one adequate. The profession is treating the raw model as a starting substrate, not a finished analyst.

This builds on work the CAS had already put in the ground. In 2025 the same working group funded research on using LLMs to process unstructured claims data, and in June 2026 it published the result: a proof-of-concept, two-stage framework that converts narrative claims information (adjuster notes, medical records, call transcripts) into structured actuarial variables for reserving, ratemaking, and claims management, built around a taxonomy of 36 candidate variables. The framework separates document-level extraction from claim-level synthesis, keeping the model on the task it is genuinely good at (reading messy text and pulling structured fields out of it) and stopping short of asking it to perform the actuarial judgment that follows.

The CAS is not stopping there. This year it also sought proposals for a benchmark suite that tests frontier models on actuarial tasks and re-tests each new model as it is released. That is a quiet admission that any answer to this question has a short shelf life.

Newsletter continues after job posts…

👔 New Actuarial Job Opportunities For The Week

We post 50 new handpicked jobs every week that match your expertise.

Here are a few of the new jobs this week:

Interested in advertising with us? Visit our sponsor page

Where LLMs Demonstrably Fail at Actuarial Judgment

Having gone through the research literature, I think the CAS's caution is well founded. Language models fail at several things that are not incidental to actuarial work but central to it.

The first is reliable arithmetic and quantitative consistency. Benchmarks have improved markedly, and the strongest current models now score at or near the top of arithmetic tests, but the failure pattern that matters for actuaries is not the average score. It is the behavior under complexity and distribution shift. Studies using symbolic variations of standard math problems have shown model performance degrading as problems grow more complex or as superficially relevant distractor information is added, even when the problem remains perfectly solvable by algorithm. An actuarial calculation is exactly the kind of task where the right answer is algorithmically available and a plausible-looking wrong answer is unacceptable. A model that is usually right and occasionally, silently, wrong is a liability in a reserving exercise in a way it would not be in a brainstorming session.

The second is fabrication under uncertainty. The documented tendency of models, including reasoning-tuned ones, is to manufacture plausible steps toward an answer when confronted with a problem that is unsolvable or outside their competence, rather than declining to answer. Research on reliable mathematical reasoning frames the core reliability property as the ability to recognize when a question falls outside one's knowledge boundary and say so, which is precisely what current models do least well. Actuarial judgment is saturated with uncertainty that has to be named honestly: sparse data, unstable tail behavior, assumptions that cannot be verified. A tool whose instinct under uncertainty is to produce confident output rather than flag the gap is misaligned with the single most important professional habit in the discipline.

The third is faithfulness of reasoning. A model's stated chain of reasoning does not reliably describe the process by which it reached its answer. For actuarial work, where the memorandum documenting how a number was derived is often as important as the number, a system that can produce a convincing rationale disconnected from its actual computation is a governance problem, not just a technical one. The explanation an appointed actuary signs has to be true, and a model's narrative of its own reasoning cannot be taken as that.

None of these are failures of politeness or polish that a better prompt fixes. They sit at the intersection of quantitative reliability, honest treatment of uncertainty, and auditable reasoning, which is roughly a definition of what actuarial judgment is for.

What I See in My Own Practice

My own experience lines up with the research more closely than I expected. The wins are real, and they sit almost exactly where the CAS says they should.

The clearest one is mechanical. Every month I used to build pivot tables and copy results across a chain of workbooks to produce a client's liability cash-flow files. An AI workflow does that now. It saves me hours, and checking it is easy because I know exactly what the right answer looks like.

The more interesting case came closer to real reasoning. A client needed missing growth factors added to a table in its policy administration system, and I first had to work out the formula behind the factors already there. Working with an AI model, I reverse-engineered that formula from the existing data, and along the way it flagged a keying error that had been sitting in the table unnoticed. That was genuinely useful work. But every rate the formula implied was tied back to the company's own product history, which its operations team confirmed. The model found the pattern. The accountability stayed with me.

And not everything should be an AI task. I recently handed some of my remaining monthly Excel processes to a contractor to automate with Power Query and VBA rather than AI. For a process that has to produce the same answer every month, a deterministic tool is the better fit. A system that is right 99 times and confidently wrong on the hundredth is the wrong tool for a reconciliation.

The line I hold is simple. If I am going to sign it, I need to be able to check it against something I trust independently of the model. Where I can do that, AI is a force multiplier. Where I cannot, it stays out of the work.

What This Means for the Profession

Put the RFP and the failure modes side by side and a coherent picture emerges, one more interesting than replacement. The parts of the actuarial workflow closest to language (reading unstructured claims text, extracting and structuring information, drafting, summarizing) are where models already add real value, which is exactly where the CAS's funded claims paper operates. The parts closest to quantitative reliability under uncertainty and auditable judgment are where models fail in ways that matter, and where the profession is deliberately not handing over the wheel. The 2026 RFP is best understood as an attempt to work out how far the boundary between those two zones can be moved through engineering, and to produce systems a practitioner can trust rather than demonstrations that a model can be prompted.

For working actuaries, the practical reading is straightforward. The immediate, low-regret gains are in the unstructured-data and drafting layer, where a model turns narrative into structured input under human review. The core reasoning and sign-off stays human because the documented failure modes land precisely on the properties that reasoning requires. And the space in between, the domain-adapted, constrained, retrieval-grounded systems the CAS is now funding, is where the genuinely open questions sit. Whether that middle ground yields tools an actuary can rely on, or mostly confirms where the human has to remain, is unsettled. That is why it is a research program and not a headline, and it is a far better question than the one about replacement.

My Take

Stop asking whether AI can replace actuarial judgment. Ask which of your tasks are language problems and which are judgment problems. Hand the first kind to the machine, with review. Keep the second kind. And keep watching the middle, because that boundary is moving faster than most of our processes are.

I would like to hear where you have drawn the line in your own work. Hit reply and tell me what you trust AI with, and what you never will.

Sources:

  • CAS Artificial Intelligence Working Group

  • 2026 Request for Proposals on adapting LLMs for specialized P&C actuarial reasoning (casact.org);

  • CAS, "New CAS Research Explores LLM Applications in Claims Analysis" and the 2025 RFP on unstructured claims data (casact.org);

  • CAS, 2026 Request for Proposals on evaluating LLMs for actuarial perception tasks (casact.org); Lieberthal et al., "Leveraging Large Language Models for Unstructured Claims Data Analysis," CAS Forum, June 2026;

  • Academic literature on LLM reasoning reliability, including work on arithmetic benchmarks, symbolic-variation degradation (GSM-Symbolic), reliable mathematical reasoning and knowledge-boundary refusal, and chain-of-thought faithfulness (2024 to 2025, arXiv and OpenReview). Model-capability findings reflect published benchmarks and are subject to rapid change.

Looking for clarity on consulting, income, or next steps?

The Independent Actuary Book Cover

Read my new book, The Independent Actuary - a practical guide for actuaries looking to build more income, leverage, and career optionality beyond the traditional path.

Buy Now At Amazon

Last week we covered The Year-End Opinion Timeline Nobody Shows the CFO.
👉 If you missed the last week’s issue, you can find it here.

💼 Sponsor Us

Get your business or product in front of thousands of engaged actuarial professional every week.

💥 AI Prompt Of The Week

Exam Study Coach

Turns AI into an adaptive study partner for actuarial exams. Quizzes you progressively and pinpoints where your reasoning is shaky.

The Prompt:

❝

I’m studying for [exam]. Quiz me on [topic] with progressively harder questions. After each answer, tell me what I got right, where my reasoning was shaky, and give me one harder follow-up.

🌟 That’s A Wrap For Today!

We’d love your thoughts on today’s newsletter to make My Actuary Weekly even better. Let us know below:

Login or Subscribe to participate