Featuring three FlexFest perspectives on AI, accountability, and the human judgment layer
For fifteen years, the modernization of legal work pointed in a single direction: toward standardization, toward repeatability, toward taking a domain full of variance and making it behave predictably. The most powerful tool to arrive in the field since then does the opposite. Understanding why starts with two words that get used loosely and mean very specific things: deterministic and probabilistic.
A deterministic system produces the same output every time it receives the same input. Feed it the same facts and it returns the same answer, reliably, on the first try and the thousandth. This is the logic of a spreadsheet formula, a docketing rule that calculates a filing deadline, a conflicts check that either finds a match or does not. It is also the logic that has governed legal technology for most of the last two decades. The NVCA model documents, the One NDA project, the spread of playbooks and clause libraries and matter templates were all efforts to make legal work behave deterministically, to force a domain full of variance toward repeatable, predictable outputs.
A probabilistic system works differently. Given the same input, it generates an output by predicting what is most likely to come next, which means the same prompt can yield different answers on different occasions. This is how generative AI works, and it is not a defect to be patched. It is the mechanism. The variability is inseparable from the capability that makes the tool useful in the first place.
That is the collision the profession is now absorbing. Lawyers have spent a generation trying to make their work more deterministic, and the field’s most powerful new tool is fundamentally probabilistic. The two impulses point in opposite directions, and reconciling them is the real work in front of every legal function right now.
That tension is exactly what Priori CEO and Co-Founder Basha Rubin put on the table at FlexFest, Priori’s inaugural virtual summit on the future of legal work, during the session on AI-native services and the end of scale as an advantage.
“All of these AI tools are probabilistic rather than deterministic technologies,” Rubin said. “We’ve all had the experience of asking AI a question and getting different answers. Ask it five times, get five different answers, and that is a feature, not a bug of how the technology works.”
She then drew the line back to where the profession had been heading. “In many ways it is anathema to the move within legal tech over the preceding 10 or 15 years, which was towards standardization. It’s sort of a pull in a different direction.”
The question that follows is the one every GC, CLO, and head of legal ops is now living with. If your modernization thesis was built on driving variance out of legal work, what do you do with a technology whose defining trait is that it reintroduces variance by design? Three FlexFest speakers answered from different angles, and together they add up to something closer to an operating framework than a debate.
Watch the clip here:
Mark Pike: Ground the model, keep the human
Mark Pike, Associate General Counsel at Anthropic, took the question to the tooling layer. Anthropic tells users plainly that its models can be wrong. “For legal work, that’s a useful disclaimer,” Pike said. “We encourage people to keep a human in the loop and ensure that for these high-risk use cases, you do have a lawyer review things before submitting.”
The mitigation Anthropic has built toward with Claude for Legal is grounding, anchoring outputs in primary and authoritative sources so that a given answer traces back to the law it rests on rather than to the model’s general sense of how legal text tends to sound. Pike likened it to giving the model a calculator to do math. A calculator does not make a person better at reasoning, it makes the arithmetic step reliable so the person can spend judgment where judgment belongs.
The more provocative part of his remarks cut against the anxiety itself. Law, Pike noted, is not deterministic to begin with. “I don’t know if there is always a correct answer in law. That’s the exciting opportunity for human judgment.” The standardization project of the last decade never actually removed indeterminacy from legal work. It contained indeterminacy inside standardized forms and processes, which is a different thing entirely. Much of the discomfort with probabilistic tools is really discomfort at being reminded that the underlying material was never as deterministic as the process made it appear.
The mature model is both, and the failures are failures of judgment
Zack Shapiro, Founder and Managing Partner at Rains LLP, an AI-native firm built on Claude, refused the framing as a choice at all.
“The mature version of all this will have both generative and deterministic elements,” Shapiro said. The dividing line is functional, not ideological. Cap tables, share counts, and citation checks against the Federal Register are deterministic tasks with correct answers, and they belong to deterministic code, not to a language model asked nicely to be accurate. Drafting contract language, capturing what a client actually wants, and expressing a complicated position in prose are generative tasks, and that is where probabilistic tools earn their place. A firm running its cap table math through a language model has made a category error, and so has a firm expecting a rules engine to draft a nuanced indemnity.
Shapiro also pushed back on the narrative that has dominated the discourse, the steady stream of stories about lawyers filing hallucinated citations before federal judges. His read reassigns the fault. “That’s not a failure of the technology. That’s a failure of judgment.” The tool produced a plausible-looking output, which is precisely what a probabilistic tool does. A licensed professional then filed it without checking, which is precisely what professional responsibility exists to prevent. On his reading, the failure was never the hallucination but the missing review step that every speaker here treats as non-negotiable.
The operational translation for legal ops follows from that. A large share of the AI risk a legal function faces is not model risk, it is process risk, the risk that a workflow lets unreviewed probabilistic output reach a place where it carries consequences. That is a controllable variable, and controlling it sits squarely within the mandate legal ops already holds.
Knowing which mode a task requires, before you deploy into it
Sabastian Niles, President, Chief Legal Officer, and Corporate Secretary at Salesforce, took the framework up to the decision a legal leader actually owns.
“When you’re deploying artificial intelligence into serious workflows, that blend, and your clarity on when do you need deterministic approaches and when are you willing to have probabilistic ones, matters,” Niles said.
The operative word is clarity. Niles is not stating a preference for one mode over the other, he is describing a diagnostic responsibility that belongs to leadership. Before a tool enters a workflow, someone accountable has to have decided which parts of that workflow can tolerate a probabilistic output and which parts require a deterministic one, and that decision cannot be handed to the tool or the vendor. It is the legal function’s judgment to make and to own.
Niles also named the deeper shift underneath all of this. AI moves lawyers away from a world of rules and toward a world of standards, from crisp yes-or-no determinations toward broader frameworks that call for more human judgment rather than less. For the highest-stakes work, closing a negotiation or arbitrating a dispute, he was direct that neither the tools nor the profession are ready to hand full authority to an AI agent. What he described instead is a hybrid in which human accountability and AI checking iterate against each other, each catching what the other misses.
The human judgment layer is doing real work
The through line across all three is that the strongest legal teams are not choosing between deterministic and probabilistic tools. They are designing workflows that route each task to the mode that fits it, and they are building an explicit human judgment layer that decides, task by task, which mode to trust.
That layer carries more weight than the industry tends to admit, and it is worth being honest about why. Niles was candid that humans are fallible, get tired, and become overwhelmed, which is exactly why the answer is not human oversight alone but a genuine pairing in which person and system check each other. Shapiro described techniques where different frontier models cross-examine one another’s outputs, adding redundancy at the machine layer. Anthropic’s contribution, in Pike’s framing, is to ground the models in the authoritative tools lawyers already rely on, the calculator for the math.
For a GC, CLO, or head of legal ops, the takeaway is not that AI is trustworthy or dangerous. It is that trustworthiness is a property of the workflow you build, not of the model you buy. The deterministic and probabilistic pieces are both available. The advantage goes to the teams deliberate about which is which, and treating the judgment layer between them as something to be designed rather than assumed.
FlexFest sessions are now available on demand. Watch here.