Procura
A case study by Sandip Sarkar, Lead Designer
Agents prepare.
People decide.
Procura replaces a chain of emails, spreadsheets and ERP screens with one workspace where AI agents gather quotes, check policy and negotiate, and accountable people make every commitment.
The core problem
Buying a fleet of laptops at a large company should not take a project manager. Today the request arrives as an email. A specialist chases quotes by hand, checks the purchase against a policy PDF and forwards a thread to approvers who have no context. Finance finds the budget problem late. Six weeks later nobody can say why a vendor was chosen.
That is slow, and it is also expensive in ways that do not show up on the invoice: rework, off-contract buying, weak audit trails and approvers who sign without reading.
The solution
A guided workspace for every request. The requester describes the need in a sentence. Agents compare vendors against policy, recommend one and negotiate. Then the work moves through six steps, selection, price, approval, purchase order and delivery, and each one closes with an explicit human action. Nine role-based views show the same record through each person's job.
The design rule that shaped everything: an agent may prepare any decision, but it may never make one.
What we measured
Measured on the working prototype with an automated audit and scripted end-to-end tests. Method and detail are in section 7.
Why enterprise procurement breaks
Four failures compound each other, and most redesigns fix only one.
- Fragmented systems
- Requests live in email, quotes in spreadsheets, contracts in a repository and money in the ERP. No one screen holds the whole deal, so every handoff re-keys data and loses context.
- Manual approvals
- Approvers get a forwarded message and must reconstruct the case. Some approve blindly to clear the queue. Others stall it.
- No shared visibility
- The requester cannot see where the request is. Finance cannot see what is coming. So both ask procurement, and procurement becomes a status desk.
- Rogue spend
- When the official path is slower than a company card, people use the card. This is a design failure, not a discipline failure, and it is the one that costs the most.
The shape of the data
The real annual budget and year-to-date spend figures seeded across the 12 departments the prototype models — the operating context every policy and threshold in this case study was designed against.
Bar length is each department's annual budget relative to the largest (IT Infrastructure); the filled portion is spend to date. $420M total budget across 12 departments and 1,200 suppliers.
Strategic objectives
The business wanted governance. Employees wanted speed. We refused to trade one for the other and set three objectives that let each support the other.
| Objective | The tension | How we resolve it |
|---|---|---|
| Compliance by construction | Policy checks slow people down | The system runs the checks while the request is being built, so the compliant path is also the fast one. |
| Speed by default | Faster often means fewer controls | Agents do the legwork. People only spend time on judgement: selecting, accepting a price, approving. |
| Every decision attributable | Audit trails add admin | The record is a by-product of working. There is no separate logging step to forget. |
Constraints and edge cases
Each constraint changed a design decision. They were not footnotes.
| Constraint | What it forced | Design response | Status |
|---|---|---|---|
| Multi-tier approval hierarchies | Routing cannot depend on who asks | Policy routes by amount and category. Under $5,000 is automatic. Above $10,000 for IT hardware needs the Procurement Manager, then the Finance Manager, in that order. Nobody sees their own request in their queue. | Built |
| Regional compliance laws | Rules differ by country and change | Policy is configuration with region as an attribute, not a forked flow. Admins change thresholds and autonomy levels without a release. | Specified |
| Legacy SAP and Oracle integrations | Budget data is delayed and sometimes wrong | Treat budget as a reservation, not a fact. Show an "as of" time and a clear conflict state. See section 5. | Specified |
| Legacy user inertia | People will not learn a new tool for a task they do twice a year | The entry point is a sentence, the way they already ask. The workspace shows the structured brief beside the chat so nothing is hidden. | Built |
How established players approach this
Three products define the current landscape, and each is making a different bet on where control should sit. This is a landscape scan from public product and analyst coverage, not a formal vendor evaluation.
| Vendor | Where the bet is placed | Where Procura differs |
|---|---|---|
| SAP Ariba | An AI-native rebuild (“next-gen Ariba”) that embeds Joule agents across sourcing, contracts and invoicing, with agentic intake becoming a standard front door. Its advantage is depth: decades of transaction data inside one system of record. | The AI layer sits on top of an existing enterprise suite. Procura is built request-first, with the deal model and the approval gate as the foundation, not an addition. |
| Coupa | Leans into autonomy by design. Its Navi agents run sourcing events, onboard suppliers and process payments, with the company reporting requisition cycle times cut by up to 50%. The stated goal is to move people from running the process to directing outcomes. | Procura treats acceptance as a step, not a checkpoint: the requester explicitly accepts a negotiated price and the approver explicitly signs, both logged by name, before anything moves on. |
| Zip | An intake-to-procure specialist most enterprises add in front of an existing stack (it integrates with SAP Ariba, Coupa, NetSuite and others). Its strength is the employee-facing intake screen: guided, policy-aware, and AI-assisted. | Procura carries that same guided-intake idea through the whole lifecycle, not just the request screen, so the same clarity applies at negotiation, approval and delivery. |
Most of the category is racing toward broader autonomy. Procura's bet is that the durable trust problem in enterprise procurement is not speed, it is attribution: knowing exactly which person said yes, every time an agent's work turns into a commitment.
Research and reframing the problem
We started by asking how to make requisitions faster. We ended up asking how to make each decision easier to trust.
The process
We combined process and policy analysis, structured design reviews, a heuristic evaluation, an accessibility audit and six scripted end-to-end walkthroughs to shape and validate the product.
Design thinking process
The same work, read through the five-stage model before the Double Diamond breaks it down further.
Empathize
Process and policy analysis, and living inside the legacy workflow's failure points.
Define
Clustered the pain into four themes and reframed them as three design objectives.
Ideate
Three approval concepts, and the ERP and async-queue trade-offs, worked through side by side.
Prototype
A 102-screen working build across 9 roles, with a real deal model behind every screen.
Test
A heuristic evaluation, an accessibility audit, and six scripted end-to-end walkthroughs.
Discover
Process and policy analysis: where a purchase actually gets stuck between request and delivery, and where thresholds and approval chains apply.
Define
Reframed four failures into three objectives, and wrote the personas' jobs, fears and what each needed from the product.
Develop
Explored three approval concepts, worked through the ERP and async-queue trade-offs, and built the deal model iteratively across 9 roles.
Deliver
A heuristic evaluation, an accessibility audit, and six end-to-end walkthrough scripts, shipping a 102-screen working prototype.
Methods, mapped to both models above
| Method | What it answers | Status |
|---|---|---|
| Process and policy analysis | How a purchase moves from request to delivery, and where thresholds and approval chains apply | Done. Produced the six-step deal model and the routing rules. |
| Design reviews with the project lead | Where the flow let someone bypass a control or left a decision unmade | Done. Surfaced the self-approval gap and the negotiation-to-approval jump. |
| Heuristic evaluation against Nielsen's ten principles | Usability gaps in the working prototype | Done. Fixed: submit buttons enabled on empty input, quantity errors shown as disappearing toasts, and no undo for a mistaken change. |
| Scripted end-to-end walkthroughs | Whether every role can finish its part, and whether numbers agree across screens | Done. Six walkthrough scripts run on every change. |
| Automated accessibility audit | Whether any screen or flow fails WCAG 2.1 A or AA | Done on 40 unique screens and 12 flow states. Results in section 7. |
Three primary personas, three different jobs
Nine roles use Procura. Three groups carry the main design tensions, and the other six reuse the same model. The goals and fears below come from process analysis and reviews of how each role interacts with the request.
- Job to be done
- Get the right equipment on time without learning procurement.
- What they fear
- Doing it wrong, then waiting without knowing why.
- What we gave them
- A sentence to start, a visible next step, and a "Waiting on you" list that says what needs them and nothing else.
- Job to be done
- Keep every purchase compliant and get good prices.
- What they fear
- Being blamed for a policy breach they could not see.
- What we gave them
- Automatic policy checks, vendor scoring, and agents that negotiate but wait for approval on anything external.
- Job to be done
- Decide quickly with enough context to stand behind it.
- What they fear
- Signing something they did not understand, or becoming the bottleneck.
- What we gave them
- A decision card with the agent's recommendation, risk and evidence, and a send-back that tells the requester exactly what to fix.
Composites built from process analysis, policy documents and design review, using the names the product itself gives these roles — not from interview transcripts.
Empathy maps: what each role carries into the request
Built the same way as the personas above — from process analysis and each role's documented fears, not verbatim interview transcripts from a specific person.
Jordan Bennett · Requisitioner
Says
- "I just need to know what to buy and get it approved."
- "Why is this taking six weeks?"
- "Can someone just tell me where this is stuck?"
Thinks
- Am I going to get blamed if I pick the wrong vendor?
- There has to be a faster way to do this than email.
- I hope I filled in enough detail that nobody sends this back.
Does
- Emails procurement with a rough description and waits.
- Pings Alexis directly to check on status instead of trusting a tracker.
- Uses a company card for anything small enough to avoid the process.
Feels
- Uncertain about whether he is doing it "right."
- Mildly anxious during the silent stretch between sending and hearing back.
- Frustrated that a routine purchase needs this much of his attention.
Alexis Ferreira · Procurement Specialist
Says
- "I need three quotes before I can even start comparing."
- "If I miss a policy check, it's my name on it."
- "Let me confirm this against the contract terms before I move it forward."
Thinks
- Is this actually the best price, or just the easiest quote to find?
- I hope this negotiation doesn't run past the delivery deadline.
- I need to know this is compliant before I put my name behind it.
Does
- Checks every quote against the category's policy threshold by hand.
- Negotiates directly with vendors before bringing a recommendation forward.
- Keeps track of which suppliers actually deliver on time.
Feels
- Responsible for compliance risk she did not create.
- Confident with full evidence in hand, exposed without it.
- Protective of vendor relationships built over time.
Taylor Whitman & Nicole Park · Head of Procurement & Finance Manager
Says
- "I can't approve this if I don't understand what I'm signing."
- "Send it back — I need to know why before I approve it."
- "Why is this in my queue again with no explanation?"
Thinks
- Am I the reason this is stuck, or is it waiting on someone else?
- If this goes wrong, can anyone tell it was approved on incomplete information?
- I want to move fast without missing something real.
Does
- Reads the recommendation and risk flag before anything else on the card.
- Requests changes rather than rejecting outright, when something is fixable.
- Checks who else has already signed before deciding.
Feels
- Accountable for a decision built on evidence someone else assembled.
- Wary of becoming either a rubber stamp or the bottleneck.
- Relieved when the evidence is complete enough to decide in one pass.
Affinity map: grouping what we found
Themes clustered from policy documents, the legacy workflow, and the design reviews that ran through the build — not a facilitated workshop.
Fragmented information
Approval friction
Invisible risk
No accountability trail
Surfacing the core pain points
Pulled from the affinity map above and ranked by how much damage each one does — a design judgment, not a measured severity score.
Features that eliminate the pain
Every pain point above maps to a specific, shipped part of the product — not a general promise.
What we learned
Each insight led to a design commitment.
| Insight | Design commitment |
|---|---|
| Approval fatigue comes from missing context, not from volume. | Every approval is a decision card carrying the evidence, the risk and a recommendation. |
| If the requester can finish every step, the other roles become decoration. | Separation of duties is enforced in the structure, not in a warning message. |
| "Sent back" with a free-text note makes the requester guess. | Reasons are structured, and each becomes a guided task the requester can complete. |
| Decisions made in chat or email leave no trail. | Every action writes to one timeline, attributed to a named person or a named agent. |
Journey: before and after
The baseline lane is the typical legacy flow we used as our working baseline.
One conceptual model, many lenses
The legacy process was a dozen tools. The product is one object with a lifecycle.
System complexity
Requests, quotes, negotiations, approvals, purchase orders, deliveries and invoices used to be separate records. We collapsed them into one: the deal. A deal moves through six steps, and every role sees the same deal through the lens of its job.
This had a practical payoff. When a requester cuts the order from 25 laptops to 20, the price, the approval amount, the purchase order, the supplier's portal and the delivery check all change together, because they all read from one place. Finance stops finding numbers that disagree.
Information architecture and permissions
Nine roles share one underlying set of functional areas. Each role's navigation is just the areas that role can act on — nothing is duplicated per role, and nothing here implies one role reports to another.
| Functional area | Employee | Proc. Specialist | Proc. Manager | Finance Mgr | CFO | Legal | Supplier | CPO | Admin |
|---|---|---|---|---|---|---|---|---|---|
| Overview (home, my work, approvals) | – | ||||||||
| Requests & sourcing | – | – | – | – | – | ||||
| Contracts & orders | – | – | – | – | |||||
| Finance (invoices, payments, budgets) | – | – | – | – | – | – | |||
| Analytics & intelligence | – | – | – | – | – | ||||
| Agent oversight | – | – | – | – | – | ||||
| Governance (policy, audit) | – | – | – | – | |||||
| Organization admin | – | – | – | – | – | – | – | – | |
| Supplier portal | – | – | – | – | – | – | – | – |
Areas are functional groupings for this diagram, not literal in-app section labels. "Access" means the role has at least one page in that area; it does not show which specific pages.
Procurement Manager and Admin both touch agent oversight and governance, but for different reasons: Procurement Manager also runs sourcing and sees analytics, while Admin alone controls users, roles and organization settings. They are peers with overlapping concerns, not a reporting line — the matrix makes that visible in a way a tree could not.
Navigation is role-based, and it shows only what a role can act on. A requester sees four items. A procurement manager sees sourcing, contracts, spend and risk. A supplier sees a separate portal, scoped to that supplier's data only. The organization layer sits above roles, with users, roles, integrations and policy settings, so the same model scales to multiple tenants and a configurable permission matrix.
Design system and component governance
An enterprise tool lives or dies on consistency, because people use it rarely and cannot afford to re-learn it. We built a tokenized library and, more usefully, a short set of rules for using it.
- One component, one job. Dense data tables, decision cards, step accordions, insight callouts and status badges each have a single purpose. A new variant needs a new job.
- Color means status, everywhere. A callout takes the tone of its step's badge: amber for "needs you", green for "done". We broke this once, caught it in review and fixed it.
- Same verb through the whole flow. The action "Select CDW" produces the status "CDW selected". Buttons say what will happen.
- Density without noise. Tables scroll inside their own container. Detail sits in progressive disclosure, one step open at a time.
- No decoration. Vector icons only, sentence case, and no color that is not carrying meaning.
Exploration, pivots and hard trade-offs
The approval workflow is where governance and speed collide, so it is where we explored the most.
Three concepts for approvals
We compared three concepts against the approver's core question, "can I stand behind this?", and against the risk of rubber-stamping.
| Concept | What it offered | Why we moved on or kept it |
|---|---|---|
| A. The queue | An inbox of requests with bulk approve, familiar from email | Fastest to learn and fastest to rubber-stamp. Bulk approval invites signing without reading, which is the failure we were fixing. We kept a narrow form of it for low-risk items. |
| B. Chat only | Approve by replying to the assistant | Natural and quick, but it gives an approver no surface to inspect evidence. We kept it as an accelerator for people who already know the case. |
| C. Decision cards in context | Evidence, risk and recommendation on one card, with the full workspace one click away | Won. It answers the approver's real question, which is "can I stand behind this?" |
Four things we got wrong, and fixed
The most credible part of any case study is what changed your mind. These four came from our own reviews of the working prototype.
- The requester could approve their own purchase. An early build let Jordan complete every step. We caught it in review. Fixing it meant architecture, not a warning: a two-approver sequence above the threshold, and nobody sees their own request in their queue. Once the requester could not finish everything, the other roles had a reason to exist.
- Negotiation jumped straight to approval. The agent finished and the screen moved on, so the requester never said "yes, I'm happy with this price". We added an explicit step: Accept $64,750 and continue to approval.
- The agent recommended and moved on. The button read "Show your recommendation", which implied the agent was deciding. We changed it to "Select CDW" and split the badges: AI recommended belongs to the agent, Selected belongs to the person.
- "Request changes" was a text box. Approvers wrote notes and requesters guessed. Now an approver ticks reasons, and each one becomes a task: negotiate again, consider another supplier with real offers, reduce the amount, or clarify the need.
Technical constraints: ERPs and asynchronous queues
In production, budget lives in SAP or Oracle, and it is not instant. Approvals are asynchronous. We designed around both rather than hiding them. The prototype simulates the responses.
| Situation | What the person sees | Why |
|---|---|---|
| Budget check in flight | "Checking budget" with a spinner, and the request can continue | Blocking on a slow system teaches people to route around it. |
| Budget confirmed | "Reserved, as of 09:41" with the remaining amount | A timestamp is honest about freshness. We treat budget as a reservation, not a fact. |
| Conflict after approval | A named exception to the owner with the gap and two options | Silent failure is the worst outcome. A conflict is a decision, so we route it to a person. |
| Approver is slow | "Waiting on Taylor Whitman, then Nicole Park" on the request, plus the SLA state | Visible waiting removes the status-desk work and makes delay attributable. |
The decision ledger
The rationale for the choices that mattered most.
| Decision | We chose | Over | Because |
|---|---|---|---|
| Approval surface | Evidence cards in context | A bulk-approve queue | Fatigue comes from missing context. Bulk approval treats the symptom and raises risk. |
| Agent autonomy | Act with approval for anything external | Act automatically | Negotiating and ordering commit the company. Only purchases under $5,000 run automatically. |
| Intake | A conversation that asks only for what is missing | A long form | Forms produce errors and abandonment. The brief stays visible beside the chat so nothing is hidden. |
| Send-back | Structured reasons that become tasks | A free-text note | Requesters should never have to guess what to fix. |
| Data consistency | One deal model behind every screen | Per-screen data | Finance stops trusting a tool when two screens disagree by a dollar. |
| Safety net | Rollback per step and undo on each change | Confirmation dialogs | People click through confirmations. Reversibility protects them without adding friction. |
| Status color | Tone always matches the badge | Color per component | Color only works as a signal when it means the same thing everywhere. |
The final experience
Three flows carry the product. Each one shows the rule in action.
Flow 1. Guided, smart requisitioning
The requester types what they need. If the item, date or budget is missing, the assistant asks for exactly that and assumes nothing. Once the brief is complete, agents check policy, find vendors and compare quotes, and the workspace fills in as they work.
- Real-time validation. The header shows the budget until a price exists, then the value. Changing a quantity shows the new total and saving before the person applies it, and an invalid value explains itself in place.
- Agent, then person. The agent marks its pick "AI recommended". The person selects. The agent negotiates. The person accepts the price. Two explicit human decisions before any approval.
- No dead ends. A rollback on each completed step returns to that point, and a chip suggests the next sensible action.
Micro-flow: from a sentence to an approval request
Flow 2. The smart approval hub
Approvers open one page of decision cards. Each card carries the amount, requester, vendor, risk and the agent's recommendation, so most decisions take one look. Anomalies lead the card. Examples from the prototype are a data-center contract flagged high risk for single-supplier concentration, and a supplier switch that delivers a week after the required date.
- Sequential sign-off
- For a $64,750 IT hardware order, Taylor signs first, then Nicole. The card shows the order and who is next.
- Updated after feedback
- When a requester resubmits, the approver sees a one-line summary of what changed. They never re-read the whole case.
- Batch actions
- Approve several low-risk items under the threshold in one action. Exceptions and anything flagged high risk are excluded by rule, so batching cannot hide a problem.
Flow 3. Vendor intelligence and the contract lifecycle
Specialists and executives see one scorecard per supplier: delivery, quality, pricing, responsiveness, disputes and trend. When a score moves, the page leads with the change and its cause, not the table. Contracts carry renewal dates and exceptions, and every request header carries an SLA state that turns amber before it turns red.
SLA compliance impact
Every request carries one of four real SLA states, shown on its header at all times. This is what the mechanism does, not a measured adoption number — that needs real usage data, which section 7 does not yet have.
On track
The default state. Nothing is waiting on a person for longer than expected.
At risk
Set the moment a request sits in approval, or a vendor commits to a date after the requirement.
Critical
A delivery exception is open and unresolved — the one state that interrupts rather than just informs.
Closed
Receipt is confirmed. The record stops moving and becomes the audit copy.
Before, an SLA breach surfaced only when someone asked why an order had not arrived. Now the state is visible on the request itself the moment the condition is true — including when the risk comes from a decision the requester made, like accepting a later delivery date, so the trade-off stays visible instead of disappearing into a footnote.
Accessibility and inclusivity
We audited every screen and flow with axe-core against WCAG 2.1 A and AA. The first pass found 26 violating elements: five filter dropdowns with no accessible name, a dimmed table row and a status label below the 4.5:1 contrast ratio, and a scrollable chat log the keyboard could not reach. We fixed all of them, and the re-audit found none.
- Keyboard first. A skip link, a visible focus ring, and "/" to search. Every action is reachable without a mouse.
- Color is never the only signal. Every status has a word and a shape as well as a tone.
- Errors explain themselves. Validation sits beside the field, states the rule and is announced to screen readers.
- State is announced. Toggles expose pressed state, and updates to the assistant and to hints use live regions.
- Dense without overload. Progressive disclosure keeps one step open at a time, and tables scroll inside their own region.
Impact: what we measured
We measured the product itself: what the interface does, how completely it records decisions, and how it performs.
Measured results
| Measure | Result | What changed | How we measured it |
|---|---|---|---|
| Accessibility violations on screens | 5 to 0 | Five filter dropdowns had no accessible name. | axe-core, WCAG 2.1 A and AA, 40 unique screens reached from 102 role views. |
| Accessibility violations in flows and dialogs | 21 to 0 | A dimmed table row, an unreachable chat log and a low-contrast status label. | axe-core on 12 states, including both dialogs and the supplier portal. |
| Human decisions recorded by name in the audit trail | 4 of 10 to 10 of 10 | The trail logged the agents' work but not selecting, accepting, sending or ordering. | A full walkthrough from the first sentence to confirmed receipt, then reading the Audit Trail page. |
| Human decisions in a standard purchase | 10, across four people | Requester 7, approvers 2, supplier 1. | Scripted walkthrough of a $64,750 hardware order. |
| Agent steps recorded for the same purchase | 14 | Each attributed to a named agent. | The audit trail of the same walkthrough. |
| Load time | 73 ms median, 431 KB in one file | Not a network figure. It shows the prototype carries no hidden weight. | Seven cold loads on a local machine with browser navigation timing. |
What the measuring caught
Each result above came from counting, not from looking. The audit trail is the clearest case. It looked complete in every screenshot, but counting showed that six of ten human decisions were missing. That contradicted the product's own promise that every decision is attributable, so we fixed it.
Leadership, lessons and what is next
Aligning the people who could say no
The product only works if risk, engineering and the commercial team all believe it. Each had a different worry, so I framed the work around it.
| Partner | Their concern | How I aligned them |
|---|---|---|
| Risk and compliance | "An AI should not be able to commit us." | Made the rule visible in the product: agents prepare, people decide. Separation of duties is structural and every action is attributed. |
| Engineering | "Design hands us a moving target." | Shared tokens and a component list with jobs. Six end-to-end walkthrough scripts act as a shared definition of done and catch regressions. |
| Sales and the client | "Will it look real in the demo?" | Realistic data, real vendor brands for positive scenarios, and every persona fully working end to end. |
What I would do differently
- Model the deal before drawing screens. We built a hundred screens and then made quantity and supplier dynamic, which meant rewriting well over a hundred hard-coded values. A single deal model on day one would have cost a fraction of that.
- Define roles before pages. The self-approval flaw appeared because we had screens for every persona but had not yet decided what each persona was not allowed to do. Write the permission matrix first.
- Test visually as well as in code. Headless tests passed while a layout bug sat in plain sight. Screenshots catch what scripts cannot, so they are part of every review now.
- Measure from the first build. The audit-trail gap and the accessibility failures were found by counting, not by looking. A measurement harness should exist from day one.
- Widen the review earlier. The prototype was shaped by expert review and scripted walkthroughs; broadening that circle sooner would have surfaced more edge cases before the build matured.
What is next
- Predictive replenishment. Suggest a request before anyone asks, using usage and lead-time patterns, and always as a proposal for a person to accept.
- A conversational procurement agent. Move from guided chat to a colleague that can be briefed in plain language, while keeping every commitment a human action.
- Live ERP sync and batch approval. Replace simulated budget responses with real reservations, and ship batching for low-risk items.
- Multi-tenant administration. Build the permission matrix editor and per-region policy configuration.
- Broader validation. Structured research with procurement teams across roles and regions, to sharpen the model with real-world variation.