Professional ServicesEY

Grounding a Big-Four AI programme in the firm's own knowledge.

A firm-wide knowledge and generation platform for one of the Big Four, grounded in decades of engagement libraries with the audit trail regulators expect.

Advisory boardroom session at a professional-services firm

Client

EY

Sector

Professional Services

Duration

18 months (ongoing)

Team

9 (senior AI engineers, product engineers, retrieval specialists, an advisory partner)

The client

About EY.

EY is one of the four largest professional-services firms in the world, with more than 400,000 practitioners across Assurance, Tax, Advisory, and Consulting. Every engagement, every workpaper, every methodology it publishes sits inside a compliance and confidentiality perimeter that reflects the sector's regulatory posture. When the firm decided to invest in generative AI at scale, it made a specific commitment: no shipped surface would rely on a model's opinion, and no surface would leak an engagement across a client boundary.

Context

Where EY was when we started.

EY's engagement teams sit on hundreds of thousands of prior workpapers, methodologies, regulatory interpretations, and industry playbooks. The intellectual capital is real; the retrieval is not. Practitioners rely on personal networks and Slack channels to find the paragraph they need, and the firm — like every professional-services firm — is looking at generative AI as both an opportunity and a live regulatory question.

Challenge

The problem, unvarnished.

  • Institutional knowledge spread across SharePoint, engagement-management systems, methodology repositories, and personal drives — none of them queryable in a way that respects the firm's IAM model.
  • Client confidentiality and engagement segregation are non-negotiable — retrieval must respect the permission model at query time, not at post-processing.
  • Generative outputs going to clients or regulators must carry citations back to the source paragraph; hallucination is a career risk, not just a UX one.
  • The firm needed a single AI platform pattern rather than a hundred practice-level experiments — each with their own eval, prompt library, and shadow retrieval index.
  • Time-to-value pressure from the executive: the platform had to reach a real practitioner surface in six months, not eighteen.
Approach

How we scoped and sequenced the work.

01

Permission-aware indexing from day zero

We built the ingestion layer to respect the firm's existing IAM model. A retrieval request only surfaces documents the requesting user is already allowed to read — engagement-team scope, client scope, and firm scope enforced at query time. No parallel permission universe.

02

Grounded generation with citations as a first-class output

Every generated paragraph carries a citation into the source document and the specific passage. Reviewers can jump to the excerpt in one click. Uncited claims are refused by the system, not softened by policy.

03

Evaluation as a product surface, not a data-science ritual

The AI platform team owns retrieval quality dashboards, evaluation runs, and the feedback loops that catch drift. Generation quality is measured continuously against a growing golden set of firm-approved answers — with alerts when new content violates existing evals.

04

Firm-wide platform, practice-level adaptation

One platform, multiple practices. Assurance, Tax, Advisory, and Consulting each get retrieval scopes, prompt patterns, and evaluation sets tuned to their subject matter — without forking the platform. A shared control plane; distinct data planes.

Timeline

How the engagement unfolded.

01

Months 1–2

Discovery & platform architecture

Two-week discovery with the firm's AI leadership, then a four-week architecture sprint pinning down the IAM boundary, retrieval pattern, evaluation harness, and generation guardrails. Decision made to standardise on Azure OpenAI with Azure Cognitive Search retrieval, Anthropic Claude as a secondary route for long-context passages.

02

Months 3–6

Assurance pilot

First practitioner surface shipped to a scoped Assurance team of 200 users. Retrieval quality baselined; the golden set moved from 40 hand-curated Q&A pairs to 400. First measured productivity delta on drafting workpaper support memos.

03

Months 7–12

Advisory expansion & shared control plane

Rollout to Advisory with practice-specific retrieval scopes and prompt libraries. Shared control plane stood up for AI platform-team ownership: eval dashboards, prompt version-control, incident response for retrieval regressions.

04

Months 13–18

Firm-wide platform hardening

Scaled ingestion to millions of documents. Cross-practice retrieval scoping formalised. AI platform team took ownership of eval, releases, and the golden-set roadmap. Tax and Consulting practice teams onboarded to the platform pattern.

Solution

What we shipped.

A permissioned knowledge and generation platform standing between the firm's document estate and its practitioners. Retrieval respects the IAM model; generation carries citations; evaluation is continuous. Deployed to Assurance and Advisory first, with a rollout path to Tax and Consulting following the same platform pattern.

Architecture & decisions

The choices behind the build.

01

Permission enforcement at query time, not post-filter

The retrieval layer respects IAM at query construction — the vector search never sees documents the user isn't permitted to read. Post-filtering was rejected as a leak risk during code review with the firm's security team.

02

Multi-model routing with a fallback tier

Standardised on Azure OpenAI for the majority of retrieval-and-summarise calls, with Anthropic Claude as the routing target for long-context passage synthesis. Model choice is a config file, not a code change; models are eval-gated before promotion.

03

Citations as structured metadata, not string concatenation

Every generated span carries structured citation metadata — document ID, passage offset, retrieval confidence. Reviewers get one-click deep-links; downstream systems can consume the citation graph without re-parsing the answer.

04

Golden set as the release gate

No new model, retriever, or prompt ships without passing the golden set. The golden set is owned by practice SMEs, not the AI platform team — the ones who know what a good answer looks like.

Rollout & adoption

How it landed inside the organisation.

Rollout was practice-by-practice, not user-by-user. Each practice got a two-week onboarding window during which its retrieval scope, prompt library, and evaluation set were finalised with the practice's senior leadership. Adoption was voluntary until the platform hit a measured retrieval-quality bar in that practice — then the firm's AI leadership sponsored a firm-wide announcement. No mandate ever went out; the firm's operating principle was that a platform partners will adopt on merit will land better than a platform mandated from the top.

Outcomes

The numbers that matter.

12M+

Documents under continuous index

SharePoint, methodology repos, engagement libraries, and prior workpapers — indexed with permission enforcement, refreshed nightly, with drift detection on retrieval quality.

72%

Reduction in first-draft author-hours

For controls documentation, prior-work summaries, and regulatory-response drafts, measured over the first three quarters of production against a manually-tracked baseline.

100%

Answers cite their sources

Uncited generation is a system-level refusal, not a policy request. Every reviewer can jump to the underlying passage in one click.

Tech stack

What we built it on.

Azure OpenAIAnthropic ClaudeAzure Cognitive SearchSharePointDatabricksSnowflakeTypeScriptPythonNext.js
Reflection

What we learned.

Two lessons carried. First, the evaluation surface mattered more than the model choice. Once partners could see a citation next to the answer, the internal conversation stopped being 'is this safe' and started being 'where else can we ship this'. Second, the IAM decisions we made in month one were the ones we couldn't have undone in month twelve — permission enforcement at query time was the single architectural call that made the platform legally shippable across practices.

What's next

The engagement today.

The platform is now the standard on which the firm's Assurance and Advisory AI programmes run. Tax and Consulting are onboarding through the same platform pattern. Roadmap includes a dedicated evaluation surface for regulator-facing outputs, an expansion of the citation graph into a firm-wide knowledge graph, and formalising practice-level 'AI-platform SMEs' as a shared operating role.

The retrieval quality is what changed the internal conversation about AI. Once we could show partners a paragraph with the source paragraph next to it, the question stopped being 'is this safe' and started being 'where else can we ship this'.

AI platform lead · EY (name withheld under engagement confidentiality)
Delivered by

Practices involved.

Products on top

Productised offerings involved.

Talk to the team

Bring us your professional services brief.

Book a 30-minute discovery call. A senior practitioner from the same practice that shipped this engagement will scope yours.