Docufy reimagines contracting as human–AI collaboration, where metadata extraction, clause governance, and risk analysis accelerate business decisions without compromising security. Embedded entirely within M365, it cut the cost of digitising a contract type from 3–6 months and $300–500K of engineering to a target of under 14 days. Review dropped from weeks to hours.
That was the finding the whole project turned on. Press With Docufy to see what the redesign removed.
Microsoft creates roughly 300 agreements on a slow day and 2,000+ at fiscal year end. At that volume, eight weeks of queue per agreement isn't an inconvenience — it's lost deals, expired offers, and vendors who move on.
Docufy rebuilt that lifecycle inside Microsoft 365, with Azure Purview holding governance and retention. Nothing leaves the compliance boundary, including during signing.
The failure that best explains it is template drift. Someone downloads a Word template to their desktop. Legal later changes one clause — a non-compete goes from one year to two. Now two versions exist, and agreements keep going out under the old terms, signed and binding, for as long as that copy survives on someone's laptop. Nobody did anything wrong. The system simply had no way to make the current version the one you actually use.
Underneath that: 1.5 million agreements in the repository growing 10% a year, and a stack that didn't talk to itself — a third-party CLM for lifecycle, SharePoint for storage, Word and email for the actual work. The technical build used SharePoint Premium's content AI services for extraction and classification, with Azure OpenAI behind the reasoning.
I owned the Word Add-in end to end — where agreements are written, reviewed and signed, and where every human–AI interaction in this case study lives. By the time SharePoint Premium joined, I was directing four designers across two orgs, which made consistency the real problem: Docufy had to read as one system across a Word pane and a SharePoint app, while still reading as Word. I set the design strategy and the trust standards both surfaces shared. They were adopted org-wide.
Day to day I worked directly with the PM and the engineers rather than through a handoff — sitting in backlog grooming, prototyping in React against the real component library, and settling feasibility questions in the same conversation where the design question came up.
The question was whether Microsoft could build an AI-powered contract system in-house, entirely within M365, instead of licensing a third-party CLM. I joined as design lead with one PM and three engineers, mapped the end-to-end flow with them in FigJam, and we built a working skeleton in a week. That demo reached the senior leadership team and became a funded cross-org initiative.
Eighteen months later the same idea was on stage at Ignite as a product, in partnership with SharePoint Premium. Worth watching this against the hackathon clip above — the distance between them is the design work.
Legacy Word templates become structured, metadata-rich models, with AI detecting and tagging the variable fields.
A central library of approved clauses and snippets, with deviation detected when someone edits the text.
An edit to governed language is compared to the approved clause and explained in terms of what it means, not what characters moved.
Risks, missing clauses and compliance gaps surfaced against the company playbook, each with a plain-language reason.
Review, approval and e-signature routed entirely within M365. Documents never leave the security boundary.
Status, obligations and relationships across every agreement, giving admins oversight across departments.
"AI-powered" is easy to claim and hard to justify. Three capabilities carried Docufy, each solving a problem with a real price tag. The pattern underneath all three: the model does the first pass, the human keeps the judgement, and the consequence of a change is shown before it's committed.
From research Metadata inconsistency · serves the Organizers
Digitising one contract type meant reading a 50–100 page template line by line, marking every variable by hand: $300–500K and three to six months, across 623 templates. AI proposes the fields instead.
Which moved the design problem rather than removing it. Detection isn't the hard part — review is. So every suggestion carries what a reviewer needs to decide in seconds: the text it matched, how many places it occurs, and where the replacement comes from — a governed field from legal's library, or one the model wrote.
Suggestions split into library fields and AI generated fields because those carry different risk — a field from legal's (CELA) library is already approved language, while a generated one is a proposal and nothing more. Every generated suggestion is labelled AI generated content may be incorrect, with a thumbs up or down beside it, so the reviewer is told what they're looking at and has somewhere to put the correction. Neither is decoration: they're what makes accepting forty fields a decision rather than a formality.
The other half of the job is authoring a clause deliberately — turning a passage into a governed snippet the company draws from. The AI proposes the name, the type and the variables inside it. What it doesn't decide is who may change it.
Negotiable and editable look like form filler and are the entire control model: set once by the person accountable for the clause, they decide what happens to somebody else's edit weeks later, in a different part of the product. This is where the Organizers' control lives — and what buys the Doers their speed downstream.
Lock both and the text can't drift — an attempt routes to the clause owner as a request. Allow editing but not negotiation and the deviation notice in §3 fires mid-sentence. Leave both open and nothing fires at all: correct for a commercial term that should vary per deal, and quietly dangerous for anything legal is accountable for.
From research Compliance and accountability · serves the Organizers
Every legal department has a playbook — the standing rules an agreement must satisfy. Those rules moved into the system, and the AI applies them before a reviewer opens the document. On the services agreement below that's 25 rules, 20 of which pass. Five don't, and they aren't equivalent: two are high risk — a missing limitation-of-liability clause, and a non-solicit that shouldn't be there — and three are low, like an NDA with no stated expiry.
The design argument is about the twenty that passed. Hiding them buys a tidier screen and an untrustworthy one: a reviewer who can't inspect what the AI approved has no basis for believing what it flagged. So the passes collapse to one line they can open on demand, the violations lead, and risk tier decides the order. This is where weeks of review became hours.
At the end of that capture the reviewer opens Validated rules and reads back exactly what the AI approved and why: payment terms are net 60, the supplier is whitelisted, the agreement references an active NDA. No violations on it, no demo value. It exists so that trusting the flags is a reasoned position rather than a leap — and the agreement can still be rejected outright at the bottom of the pane.
Review is rarely one person, and the order matters — finance reviewing before legal has settled terms is wasted work. So the reviewer sequence is defined up front, in the template, by the people accountable for it.
From research AI opportunities — the “why” behind a document · serves the Doers
The capability I'd defend hardest. Edit governed language and the system explains what the change means, not what characters moved — and it runs while you write, not three weeks later in review.
The canonical case: IP ownership goes from “owned solely by the Client” to “jointly owned”. Four words. A character diff renders that accurately and tells you nothing.
What a non-lawyer needs to know is that the Client just gave away exclusive ownership of everything produced under that contract, and that it can't be undone after signature without an amendment. Nothing in a diff says that.
Running it at the moment of the edit rather than at review is a decision about where control belongs. Feedback three weeks later reaches someone who has moved on, in a document they no longer remember writing; feedback mid-sentence reaches the person making the choice while they're still making it. It also means most deviations never reach legal at all.
Two weeks of contextual inquiry with 16–18 employees across HR, legal, real estate, finance and records — watching real workflows rather than asking about them, because nobody reports the waiting.
High document volume led to confusion, duplicate files, and version conflicts, especially across global teams.
Both personasLegal and finance needed strict control over permissions, version history, and audit trails.
OrganizersUsers struggled to locate documents or see how one related to others — contracts, amendments and renewals were disconnected.
DoersManual tagging caused errors that later broke compliance reports and retention policies.
OrganizersParticipants didn't ask for faster search. They wanted the why behind a document — what it obliges them to, what deviates, and what it relates to. That's a reasoning problem, not a retrieval one.
Both personasTwo users, opposite needs. Organizers — legal, procurement, compliance — want control. Doers in HR or finance want speed. That reads as a conflict until you notice they're asking at different moments: Organizers want control before anything is drafted, Doers want speed during. So control moved upstream into the template itself, and everything downstream feels fast to a Doer precisely because an Organizer already constrained it. Nearly every decision in the creation flow is that same move.
Accountable for templates, compliance, obligations and metadata accuracy across the organisation.
Non-legal employees who initiate agreements — vendors, NDAs, SOWs — without being experts in compliance language.
Most of the hard thinking happened before any interface existed — deciding what a “content block” actually was, what a template owner needed to control, and how much metadata to expose before it became noise. Three questions framed it.
How might we convert 100-page Word templates into structured, reusable components with built-in metadata and clause tracking?
How could users instantly see where AI made an assumption, or flag a clause that differed from legal's approved language?
How could the system show the connections between NDAs, renewals and amendments as a living ecosystem?
Two problems here couldn't be solved in a design file. When the tool I need doesn't exist I build it; when a design can't be judged from a picture I code it and put it in front of people. That's how I work generally, not a one-off on this project.
Answering “is this Adaptive Card buildable?” meant digging through the component library or booking an engineer — so the question went unasked and came back later as rework. I built AdaptiveLens, a Figma plugin on OpenAI and Gemini models that turns a frame into schema-valid JSON, names what won't render, fixes it, and shows you the React version right there.
The triptych above is the loop it closed — frame, schema-valid JSON, and the card as the client actually draws it, all three in sync and checkable before a line of production code was written.
Feasibility stopped being a gate at the end of design and became something you checked mid-thought, which is the only point at which it's cheap to act on. It's the same instinct behind the MCP workflow I run now — Figma, VS Code and Copilot against one live component state — and behind the agent answering questions on this site.
analysis method: ai-vision), validates the output, and names what won't render — here two unsupported Input.Text elements. Fix errors resolves them and re-renders, which is the state on the right. Telling a designer their card is invalid is a bug report; fixing it in place is a tool.It stopped being mine fairly quickly. Other designers picked it up for their own Adaptive Card work across the org, which is the outcome I actually wanted — a feasibility question that used to cost an engineer's afternoon became something any designer could answer in a few seconds, on their own, mid-design.
The second problem: Figma could show a dashboard holding twelve contracts. It could not show one holding twelve hundred — which is the only version that tells you whether the design works.
So I have a process, and I run it on everything now: generate synthetic data matching the real distribution rather than the convenient one, build the screen in React against it, then run the sessions on that running build instead of a clickable mockup.
Below is the same component under both conditions. Twelve tidy rows answer no questions. Five hundred realistic ones surface five decisions I'd otherwise have shipped without making — select any flagged row to see which.
A participant getting lost is data. A participant hitting the edge of a prototype is noise.
Participants sorted, filtered, opened the wrong agreement and backed out of it — none of which a clickable mockup can produce, because a mockup only supports the path you thought to draw. So the feedback was about the design instead of about the seams, and the questions came back real: why is this one red, where did the expired ones go, can I see just mine.
The side effect was cheaper handoff. Engineering received resolved behaviour — truncation, empty states, sort order, which column collapses first at narrow widths — rather than a mockup and a list of open questions. The closer design gets to the running thing, the fewer decisions get made by default, and defaults are where AI products quietly go wrong.
The number I'd lead with is the cost of digitising a contract type — what made intelligent contracting impossible at scale before this, and the figure Microsoft Digital quoted publicly at Ignite.
Adopted across Legal, HR, Real Estate and Finance after the Microsoft Digital pilot, and the blueprint for SharePoint Premium's AI-powered agreement solution — demoed at Microsoft Ignite 2023, design framework adopted org-wide.
Five insights came out of the foundational study. Two shipped outright, two landed partly, and one I scoped into a later phase. Closing the loop honestly matters more than claiming a clean sweep — the gaps are as informative as the wins.
Live modelPowered by GPT-4o Grounded in Sayena's case studies, not a canned bot. Billed per question — so make them good ones.