A practical guide to designing with evidence, not opinions dressed as taste. Methods describe how you discover and iterate. Standards describe the quality floor for public services. Keep those ideas separate so you do not bury teams in research theatre, or ship interfaces that fail the people who need them most.
MAP-01Design maps
MAP-01
Design & research maps
Pick the map that matches the uncertainty in front of you.
Civil Service + Universal
UX work mixes discovery methods, service design, interaction craft and public-service standards. Confusing them creates either endless research or a UI factory. Every framework below answers a different question.
Maps you will actually use
Map
Answers
Shape
Double Diamond
Are we solving the right problem before we polish a solution?
Discover · Define · Develop · Deliver
Jobs to be Done / user needs
What progress is the user hiring this for?
Needs, jobs, constraints, exclusions
Service design
What happens across channels, people and systems?
Blueprints, touchpoints, backstage
WCAG
Is the interface inclusively usable?
Perceivable · Operable · Understandable · Robust
GOV.UK Design System
What should we reuse by default in government?
Styles · components · patterns
GDS Service Standard
Is this fit for a public service?
14 points · assessment evidence
Choose in practice
Problem unclear: Double Diamond Discover before high-fidelity UI.
Multi-channel or operational reality: service blueprint before screen polish.
Stakeholder taste war: return to user needs and evidence, not louder opinions.
Public interface: WCAG AA in critique and acceptance from the first prototype.
Government UI: Design System by default; document exceptions with user evidence.
If a method never changes a design or backlog decision, stop performing it.
In DDaT terms, design roles run through interaction, user research, service design and content design specialisms. Titles vary. The accountability does not: someone must own whether the experience is usable, inclusive and grounded in real needs. If that person is "whoever has the strongest opinion in the room", you do not have design ownership.
Accountability ladder (typical digital design path)
Associate / Junior designer: craft on a contained journey with coaching.
Designer: accountable for experience quality on a product or service slice.
Head of Design / Design Director: owns the profession frame, quality bar and hiring.
What you stop and start doing as you grow
You stop doing
Being the only person who can draw a flow
Measuring your week by screens produced
Accepting late a11y as normal
You start doing
Coaching POs and BAs so research questions arrive early
Publishing the critique and research cadence
Measuring your week by decisions changed and inclusion risk reduced
Three questions to run every day
Whose need is this for? What evidence would change our mind? Who might we exclude if we ship this as designed?
If any answer is "unclear" for more than a week, that becomes the priority over the next visual polish request.
The Double Diamond separates problem space from solution space. Delivery pressure often collapses both into "just design the screens". Your job is to timebox discovery so learning is real, then converge hard enough that build can move.
Phases and red flags
Phase
Purpose
Red flag
Discover
Widen understanding of users, context, constraints
Interviews with no decision attached
Define
Converge on the problem and success measures
Problem statement rewritten to fit a favourite solution
Develop
Explore solution options through prototypes
One concept polished as if it were proven
Deliver
Ship thin slices and learn in the real service
Big-bang UI with no instrumentation or research follow-up
Name the decision each research spike must unlock before you book participants.
Pair with Product Owner so Define outputs land in backlog language, not only slides.
Keep dual-track: thin discovery beside delivery, not a six-month freeze.
In government, map Discover/Define evidence to Service Standard user-needs points early.
Good research reduces specific uncertainty. Bad research generates insights nobody asked for and nobody uses. Choose methods for the question, recruit for the risk, and put findings next to the backlog item they affect.
Method fit
Question type
Lean toward
Watch out
What is going on in context?
Interviews, contextual inquiry, diary
Leading questions that confirm the brief
Can people complete the task?
Usability testing on realistic tasks
Testing only happy-path demos
Which option performs better?
A/B or prototype comparison
Underpowered tests treated as gospel
How big is the problem?
Analytics, support, survey (carefully)
Vanity metrics without journey cuts
Template · Research decision card
Decision at stake: …
Uncertainty: …
Method & sample: …
What would change our mind: …
Owner of the decision: … · By when: …
Store findings where the team works, not only in a research graveyard folder.
Include excluded and edge users when risk is high; average users hide accessibility and equity harm.
Celebrate killed ideas publicly so research is not punished for inconvenient truth.
Service design looks end to end: user stages, staff roles, systems, policies and waiting. Interaction design without service thinking creates local optima that break operations.
Blueprint lenses
Frontstage
User actions and emotions
Channels and touchpoints
Evidence of progress
Backstage
Staff actions and hand-offs
Systems and data flows
Policies and SLAs that shape waits
Map the as-is before sketching the to-be when operational pain is unclear.
Invite operational staff into research and critique; they hold constraints designers invent around.
Design for failure and recovery paths, not only the marketing happy path.
Measure service outcomes (completion, contact demand, time-to-outcome), not only UI satisfaction.
If backstage work grows invisibly, you did not simplify the service. You relocated the mess.
WCAG 2.2 AA is the common public-sector floor. Inclusive design goes further: it asks who is excluded by assumptions about devices, language, cognition, time and trust. Pair standards with real users who use assistive technology.
Shift accessibility left
Moment
Design move
Smell
Critique
Keyboard, contrast, content structure, errors
"We'll check a11y later"
Prototype
Realistic labels, focus order, alternatives
Pretty mock that cannot be evaluated
Ready
Acceptance criteria include WCAG checks
Implied "make it accessible"
Done
Manual + automated evidence before release
Only automated scan at the end
Prefer Design System components that already carry accessible behaviour.
Test with keyboard and a screen reader on critical journeys every Sprint that ships UI.
Content design is accessibility: plain language, error text, status messages.
Log residual accessibility risk with owners; do not hide it behind "known issues".
Design systems encode researched patterns so teams stop re-solving solved problems badly. In UK government, the GOV.UK Design System is the default starting point for public-facing services.
When to use, when to extend
Situation
Default
Exception path
Common form, navigation, question patterns
GOV.UK / team system components
Only with evidence the pattern fails users
New domain interaction
Compose existing patterns first
Prototype, test, contribute back if durable
Brand expression
Within system tokens
Do not break accessibility for decoration
Document exceptions: why, evidence, revisit date.
Partner with frontend engineers on contribution and debt; a system nobody can implement is fiction.
Keep content patterns (questions, errors) as first-class, not only visual components.
Service assessment regularly asks why you diverged from the Design System. Have the answer before the assessor does.
Healthy critique is structured, timed and tied to goals and evidence. Unhealthy critique is opinion cosplay. Your job is to facilitate the former and interrupt the latter without becoming defensive of every pixel.
Critique rules of the road
Start with the goal, constraints and users. No feedback before context.
Ask for clarifying questions before opinions.
Separate must-fix (usability, a11y, policy) from taste.
End with decisions or experiments, owners and dates.
Invite silence from the highest-paid person until others have spoken (anti-HiPPO).
Template · Critique notes
Goal / user: …
What worked: …
Problems (severity): …
Evidence or assumption: …
Next experiment / change: … · Owner: …
If critique never changes the design, you are running a show-and-tell.
In UK public services, design and research evidence feed service assessments across Discovery, Alpha, Beta and Live. Scrambles happen when artefacts were never kept findable.
Keep warm
Evidence
Research repository with decisions linked
Accessibility approach and known gaps
Journey maps and service blueprints
Design System exceptions log
Handover
Intent, not only final screens
Content and error patterns
Analytics and research questions still open
What must not regress
Attach evidence cards to major backlog themes continuously.
Practice assessment questions in Sprint Reviews before the formal panel.
Partner with Delivery Manager and Product Owner so design is not the lone evidence owner.
The evidence does not change between a research playback and a board update. The altitude, the framing and the ask do. Design loses influence when it only speaks aesthetics, or when it hides bad usability news until UAT.
Audience calibration
Audience
Detail level
What they want
Team
High, flows, edge cases, a11y notes
Buildable intent and acceptance clarity
PO / BA / DM
Medium, decisions and risks
What changes in backlog and scope
SRO / board
Low, user outcomes and inclusion risk
Recommendation and a clear ask
Users / participants
Plain language, respect, feedback loop
To be heard and see what changed
Template · Research playback
Decision at stake: …
What we learned: three findings max
What we will change / not change: …
Open risk: … · Ask: …
Template · Design risk brief
User impact if unchanged: …
Options: A / B / C
Recommendation: …
Evidence: … · Decision by: …
Words to drop
Drop
Use
"Users will love it"
"Evidence so far: … Remaining risk: …"
"That's ugly"
"This fails [heuristic/WCAG/task] because…"
"Stakeholders want…"
"[Named person] asked for X. Recommendation is Y because…"
Under pressure, the instinct is to add more screens, more workshops or more polish. Diagnose first. HiPPO decisions, research theatre and late accessibility are usually symptoms of unclear decision rights, vague research questions or Missing Done criteria.
Symptom
Likely cause
First move
HiPPO wins every critique
No evidence bar; highest pay speaks first
Reset critique rules today; evidence before opinion; HiPPO speaks last
Research theatre
No decision owner or decision at stake
Write a research decision card; cancel work without one
Late accessibility panic
A11y not in critique, Ready or Done
Add WCAG checks to critique and DoD this Sprint
Design system ignored
Novelty rewarded; exceptions undocumented
Default to system; log exceptions with evidence and revisit date
Findings never change backlog
Playback without PO commitment
End playback with ordered backlog edits owned by PO
Endless Discover
Fear of deciding; no timebox
Timebox; ship a falsifying prototype slice
UI factory, no discovery
Date pressure eats learning
Protect a thin dual-track spike next Sprint
Assessment scramble
Evidence not kept warm
Create findable research and a11y index this week
Ops overwhelmed after launch
No service blueprint / backstage design
Map as-is backstage; redesign hand-offs before more UI
Taste wars in refinement
Goals and users not restated
Restart from task success and inclusion criteria
The one-minute checklist
Name the user need and exclusion risk in one sentence each.
Is the pain discovery, decision rights, accessibility, system reuse or service backstage?
Which decision unlocks the most learning or inclusion, and who owns it?
Tell the PO / SRO before they hear a distorted version.
Protect research and critique time while you sort the noise.