From fragmented MDF spend to a trusted global ROI platform.

How I led a cross-functional data and analytics program spanning 460K+ partner-marketing touches, three enterprise systems, seven currencies, and three global regions — turning fragmented marketing data into a decision platform partner leadership could actually trust. I owned the program end to end: framing the problem, decomposing it into workstreams, aligning technical and business stakeholders, driving the hard technical decisions, and turning a one-off analysis into a repeatable operating capability.

Program at a glance

An ambiguous global data problem, led into a structured program.

460K+Partner-marketing touches resolved to real accounts
3 → 1Enterprise systems unified into one attribution model
7Currencies normalized to USD, across 3 regions
~82–85%Touches auto-resolved; the rest reviewed or held, not forced
Artifact · Executive snapshotwhat this program was
My roleTechnical Program & Delivery Lead — owned the program from problem framing to operating model
Program objectiveGive partner leadership one trusted view of what every Market Development Fund (MDF) dollar actually sources
Business problemPartner-marketing investment, CRM pipeline, account identity, geography, currency, and campaign taxonomy were fragmented — so MDF ROI didn’t exist as a number anyone could trust
Scale & complexity460K+ partner-marketing touches · 3 enterprise systems · 7 currencies → USD · North America / EMEA / Asia Pacific
Core technologiesBigQuery · Python (identity resolution) · Salesforce · SAP · Tableau · SQL
Primary stakeholdersPartner global / executive leadership · regional geo leaders · regional partner operations · source-system owners
Operating cadenceMonthly refresh, match-quality validation, spend & pipeline reconciliation, and publication
Automated resolution~82–85% of touches resolved automatically; borderline records reviewed by a person, low-confidence records held as an explicit unattributed line
Primary outcomeFrom no trustworthy MDF attribution to a single monthly decision view leadership used to judge partner and regional performance and steer where MDF gets invested

The business problem

The problem was not a missing dashboard. It was fragmented truth.

Every cycle, the business funded partner-led demand generation worldwide — events, webinars, telemarketing, campaigns — through a global network of resellers and distributors. Leadership’s question was simple: what does that spend actually return, and where should the next dollar go? The data needed to answer it lived in disconnected pieces.

Investment disconnected from outcomes

MDF funding claims lived in the partner-marketing system; the pipeline they were meant to drive lived in the CRM — with no reliable way to connect one to the other.

Account identity didn’t line up

Customer and partner names in the claims were hand-typed free text — abbreviations, legal suffixes, typos, multiple languages — with no shared key to the CRM’s account master.

Global data wasn’t comparable

Spend spanned seven currencies and three regions, with years of inconsistent free-text taxonomy for activities, objectives, products, and solutions — none of it comparable as-is.

Why it mattered — the questions leadership could not answer

Until the pieces were connected, none of these had a defensible answer. The program existed to make them answerable.

The distinction that framed everything

This was never “we need a dashboard.” The technical solution existed to enable those decisions — and every design choice was judged by whether leadership could trust the answer enough to move money on it.

My role & program mandate

What I owned — and where I went hands-on

My mandate was to turn an ambiguous business need into an executable, cross-functional technical program — and to keep it trustworthy enough that leadership would act on it. I led the program; I also went hands-on in the technical core where that’s where the program decisions were won.

What I owned — the program
  • Problem framing — converting “understand MDF ROI” into a defined program.
  • Program strategy & scope — what was in, what was deferred, and why.
  • Workstream decomposition — six capabilities, sequenced by dependency.
  • Roadmap & milestone planning — the delivery sequence and refresh schedule.
  • Requirements translation — leadership questions into technical capabilities.
  • Stakeholder alignment — global heads, geo leaders, source owners, partner ops.
  • Cross-system coordination — three enterprise systems into one model.
  • Dependency management — the chain from source data to leadership decision.
  • Technical decision facilitation — driving defensible calls on the hard trade-offs.
  • Data-quality & exception strategy — confidence thresholds and the review path.
  • Validation & reconciliation — spend and pipeline reconciled each cycle.
  • Delivery cadence & backlog — a repeatable monthly operating model.
  • Executive decision support — the views leadership used to steer investment.
  • Operationalization — turning a one-off analysis into a standing capability.
Where I went hands-on

I used technical depth directly — not to be the engineer of record on everything, but to make better program decisions:

  • SQL & BigQuery — harmonization, taxonomy, spend logic.
  • Python — the identity-resolution pipeline.
  • Data models — touch → account → pipeline attribution.
  • Identity matching — cleansing, blocking, similarity scoring.
  • Data-quality investigation — tracing why numbers didn’t reconcile.
  • Reconciliation — spend and pipeline validation each cycle.
  • Prototyping — testing feasibility before committing the approach.

I went deep enough to validate assumptions, understand engineering trade-offs, investigate quality issues, and improve decisions — then led the program around them.

The message in one line

I was technically hands-on where it improved program decisions — but my primary responsibility was leading the program: framing the problem, decomposing the work, aligning the people, driving the technical decisions, and making the whole thing repeatable. I understand the technology deeply enough to challenge assumptions, evaluate trade-offs, and identify risk — not just track engineering work.

Program strategy

How one ambiguous request became an executable program

“We need to understand MDF ROI” is a sentence, not a plan. I decomposed it into a chain of technical capabilities — each one answering part of the business question and feeding the next. The visual below is deliberately a program decomposition, not just a data flow: every stage is a workstream with its own problem, deliverable, dependency, and exit condition.

Artifact · Question-to-decision decompositionbusiness question → investment decision
“Where should the next MDF dollar go?” → Source integration → Data standardization → Identity resolution → Attribution → Quality & reconciliation → Executive analytics → Investment decision
Read it this way

Each arrow is also a dependency: get standardization wrong and identity resolution degrades; get identity wrong and attribution can’t be trusted; if attribution can’t be trusted, leadership won’t move money — and the program hasn’t delivered its value. The chain is the reason this was run as a program, not a script.

Stakeholder alignment

Aligning the people behind the data

This was as much an alignment problem as a technical one. The same numbers meant different things to different groups — and each would only trust the platform if it answered their question. Part of my job was making one model serve all of them without bending the data to any of them.

Artifact · Stakeholder mapconcern → information needed → decision enabled
StakeholderWhat they cared aboutInformation they neededDecision it enabled
Global / executive leadershipTrusted ROI and how regions compareMDF ROI, partner performance, regional trendsWhere to invest MDF next, globally
Regional / geo leadersTheir own partners and activitiesRegional pipeline, partner performance, exceptionsLocal investment and partner management
Source / technical ownersStable definitions and reproducibilitySource mappings, identity logic, quality thresholds, traceabilityA reliable, repeatable data foundation
Regional partner operationsDefensible attribution behind payoutsThe exception queue and match-confidence scoresVerify borderline matches before they count

The principle that earned trust

Partner leaders had been burned by numbers that didn’t reconcile before. So trust was not created by pretending the data was perfect. It was created by making uncertainty visible and manageable — surfacing match confidence, showing the unattributed slice honestly, and making the same number mean the same thing in every region’s view.

Workstream decomposition

Six workstreams, one program

Each workstream solved a distinct problem, produced a specific deliverable, carried a dependency I had to manage, forced a decision, and had an exit condition before the next could safely start. Together they laddered up to the business capability leadership actually wanted.

01Source & intakeMarketing, CRM, ERP into one place
02HarmonizationGeo, currency, taxonomy, spend logic
03Identity resolutionTouch → real account
04AttributionAccount → pipeline, by rule
05AnalyticsExecutive & regional reporting
06OperationalizeRepeatable monthly refresh
01 · Source data & intake · Problem

The story was split across the partner-marketing system, Salesforce, and SAP — each with its own account identifiers.

Deliverable

One place (BigQuery) holding the three sources with their raw fields retained.

Dependency · Decision

Source-owner access. Decided: land raw first, transform downstream — so source data is replayable when logic changes.

Exit condition → capability

All three sources loading reliably → a single foundation to build on.

02 · Data harmonization · Problem

Free-text geography, seven currencies, and years of inconsistent activity/product taxonomy made nothing comparable.

Deliverable

A clean geo → country → area hierarchy, currency normalized to USD, and a stable taxonomy for activities, objectives, and solutions.

Dependency · Decision

Agreed definitions with source owners. Decided: spend follows claim lifecycle — in-flight claims use the requested amount, closed claims use the paid amount.

Exit condition → capability

Every record comparable like-for-like → global numbers that add up.

03 · Identity resolution · Problem

Hand-typed, multilingual company names had no key to the CRM. An exact-string join would match almost nothing.

Deliverable

A Python entity-resolution pipeline resolving each touch to exactly one canonical account, with a confidence score.

Dependency · Decision

Depends on clean harmonized input. Decided: country blocking before fuzzy comparison, and confidence thresholds instead of forcing every match.

Exit condition → capability

~82–85% auto-resolved, the rest routed or held → touches tied to real accounts, defensibly.

04 · Attribution & business rules · Problem

Even with matches, leadership needed a credit rule they could defend in a review — not a black box.

Deliverable

A rules-based, account-based attribution model: touch → account → that account’s sourced and influenced pipeline.

Dependency · Decision

Depends on trusted matches. Decided: explainable rules over a statistical (Markov / Shapley) model, because every credited dollar has to trace to a specific touch and account.

Exit condition → capability

Every credited dollar traceable → MDF ROI and cost-per-opportunity leaders can stand behind.

05 · Analytics & executive reporting · Problem

A correct model is useless if leaders can’t read the answer to their own question.

Deliverable

A Tableau suite: MDF ROI, sourced pipeline, cost-per-opportunity, by partner, geo, activity, and solution — with match confidence surfaced.

Dependency · Decision

Depends on the attribution model. Decided: tailor the read per audience — global heads see ROI and reallocation; geo leads see their own partners.

Exit condition → capability

Each stakeholder answered in their own view → decisions leadership could actually make.

06 · Operationalization · Problem

A number that’s true once and stale next month rebuilds no trust. It had to land reliably every cycle.

Deliverable

A repeatable monthly refresh with match-quality validation, spend and pipeline reconciliation, exception review, and publication.

Dependency · Decision

Depends on all of the above holding. Decided: run it as a standing operating cycle with an enhancement backlog, not a one-off.

Exit condition → capability

Same trusted number, every month → a durable decision capability, not a project artifact.

Technical architecture

The architecture, explained through the program

The architecture came after the problem, mandate, and workstreams — it exists to serve them. It’s deliberately layered: unify and harmonize the raw data, resolve every touch to a real account, then attribute and report. Here it is source to dashboard — and why it’s shaped this way.

Artifact · Data platform architecturesource → attribution model → dashboard
MARKETING PLATFORM Touch tracker + funding claims SALESFORCE Account master SAP Account master BIGQUERY + PYTHON ENGINE 1 · Cleanse & harmonize (SQL) 2 · Translate non-Latin names 3 · Fuzzy identity matching 4 · Currency → USD normalize ✓ 460K+ TOUCHES RESOLVED ATTRIBUTION MODEL Touch → account → pipeline TABLEAU · MDF ROI SUITE MDF ROI · sourced pipeline Cost / opportunity · partner & geo Monthly & quarterly trends
Why it’s structured this way
  • Raw retained on landing — so transformation logic can change and be replayed without re-pulling sources.
  • Harmonization centralized — one taxonomy and currency logic, so every downstream number is consistent.
  • Identity resolution isolated as its own stage — it’s the make-or-break step, and isolating it made it testable and tunable.
  • Attribution kept explainable — a rules layer on top of trusted matches, not a black box, so leaders can defend it.
The failure points that mattered
  • Harmonization errors silently degrade match quality — the most dangerous failure because it’s invisible.
  • A wrong match poisons attribution for that account — worse than a missing one.
  • Currency or taxonomy drift breaks cross-region comparability — the whole point of the platform.
  • Each is why the stage has its own quality gate, covered in Risk & quality.

The unglamorous layer — parsing a delimited partner string into a clean geo hierarchy, writing spend logic tied to each claim’s lifecycle, collapsing years of free-text into a stable taxonomy — is what every downstream number depends on. The full mechanics are in the technical deep dive.

Decisions that shaped the program

The trade-offs I drove — and why

These are the calls that decided whether leadership would trust the platform. Each was a technical decision with a business consequence — and the point of showing them is not sophistication for its own sake, but that I understood the choices deeply enough to make them defensible on accuracy, cost, explainability, risk, and stakeholder trust.

Decision 1 · Context

Leadership had to defend credited ROI in partner reviews. The attribution model had to be trusted and explained.

Options

An explainable rules-based, account-based model — vs. a more “sophisticated” statistical model (Markov / Shapley).

Decision & why

Rules-based, explainable. Every credited dollar traces to a specific touch and account. A black box would have looked smarter and been indefensible in a QBR.

Risk / consequence

Trades statistical nuance for adoption. The right trade: an attribution number leaders won’t stand behind is worthless, however elegant.

Decision 2 · Context

Partner payouts and investment decisions ride on these matches. A wrong match is more damaging than a missing one.

Options

Force every touch to a best-guess account — vs. confidence thresholds with human review and an unresolved bucket.

Decision & why

Thresholds + human-in-the-loop. High confidence auto-resolves; borderline goes to regional partner ops; low confidence is held. A person decides the ambiguous cases, not an algorithm guessing.

Risk / consequence

Accepts a review workload and a visible unattributed slice — in exchange for a number leadership can trust with money on it.

Decision 3 · Context

Comparing every touch against every account is accurate in theory and prohibitively expensive at 460K+ records — and generates cross-country false matches.

Options

Compare every possible pair — vs. country/geographic blocking before fuzzy comparison.

Decision & why

Block by country first. Only compare candidates within the same country. Cuts the comparison space by orders of magnitude — better accuracy and lower compute cost.

Risk / consequence

Assumes country is reliable enough to block on — handled in harmonization, so the assumption is controlled rather than blind.

Decision 4 · Context

Non-Latin-script markets (Japan, Korea, Taiwan, Israel) need translation to compare — but translation is the most expensive step in the pipeline.

Options

Translate/transliterate everything — vs. targeted handling only for the records that actually need it.

Decision & why

Targeted path. Isolate and translate only the small share of non-Latin records, on a separate path — keeping the monthly refresh fast and cheap without losing coverage.

Risk / consequence

Adds a branch to maintain — worth it: it kept a recurring cost-and-time driver small at scale.

Decision 5 · Context

Fuzzy matching on multilingual free text never reaches 100%. A tidy dashboard could hide that; a trustworthy one can’t.

Options

Force 100% attribution for a clean-looking result — vs. an explicit “Unattributed Partner Spend” line.

Decision & why

Keep an honest unattributed bucket. Leaders see the true size of what we couldn’t attribute, so they trust everything above it. An honest gap protects the credibility of the whole platform.

Risk / consequence

A less “perfect” headline — deliberately. Credibility mattered more than a falsely clean 100%.

Dependency management

Managing the dependency chain

This program was a chain, and the chain was only as strong as its weakest link. Managing those interdependencies — in business language, not just technical — was the core of the delivery job.

Artifact · Dependency chainsource data → leadership decision
Source data→ Standardization→ Geo / currency / taxonomy→ Identity resolution→ CRM account connection→ Pipeline connection→ Attribution→ Reconciliation→ Reporting→ Leadership decision

The chain in business language

If harmonization is wrong, matching quality falls. If account matching is wrong, attribution becomes unreliable. If attribution is unreliable, the analytics can’t be trusted. And if leadership doesn’t trust the analytics, the technical program hasn’t delivered its intended business value — no matter how well the pipeline runs.

Artifact · Dependency registerwhy it matters → downstream impact → control
DependencyWhy it mattersDownstream impact if it failsProgram riskControl
Source availabilityNo data, no programNothing refreshesCycle slipsMonthly intake check before the run
HarmonizationEverything compares on itSilent match & comparison errorsInvisible degradationTaxonomy & currency validation
Identity resolutionThe make-or-break joinAttribution unreliableFalse positivesConfidence thresholds + human review
CRM / pipeline connectionTurns matches into valueNo pipeline to creditStale account linksReconciliation against CRM each cycle
ReconciliationProves the numbersDistrusted outputUnexplained varianceSpend & pipeline reconciled before publish

Risk, quality & program controls

Where a wrong number was the real risk

The central risk here wasn’t a failed pipeline run — it was a confident wrong answer. Leadership was moving investment on this output, so I treated identity resolution as program risk, not just an engineering task.

The principle that set the quality bar

A false positive was more dangerous than an unresolved record — because a wrong match quietly credits the wrong partner and skews an investment decision, while an unresolved record is at least visibly unknown. So the whole control model was built to avoid confident wrong answers, not just to maximize match rate.

The confidence control model

High confidence

Automatic resolution. Clear, high-scoring matches resolve without a human — roughly 82–85% of records.

Medium confidence

Human-in-the-loop review. Records scoring ~0.60–0.79 route to regional partner operations to verify — a person, not a guess.

Low confidence

Held as unattributed. Below ~0.60, tagged “Unattributed Partner Spend” rather than force-matched — visible, not hidden.

Artifact · Risk & control matrixrisk → control → fallback
RiskBusiness impactPreventive controlDetectionException / fallback
False-positive matchWrong partner credited; skewed investmentConfidence thresholds; harmonic-mean scoringMatch-confidence score in the outputHuman review below threshold
Missing identifiersRecords can’t resolveMultiple signals (name, domain, tax/VAT ID)Unresolved-rate trackingHeld as unattributed
Multilingual recordsAPAC touches under-attributedTargeted translation pathPer-region resolution ratesReview queue for the residue
Taxonomy / currency driftRegions stop being comparableStable taxonomy; currency → USDCross-region reconciliationCorrect & re-publish
Silent quality degradationDistrust builds unnoticedMonthly QA & reconciliation gatesVariance vs. prior cycleInvestigate before publish
Stakeholder distrustPlatform ignoredSurface confidence & the unattributed lineStakeholder feedbackMake the seams visible, not hidden
Precise language

The ~82–85% figure is an automated resolution rate — the share of records that matched confidently without human review — not a validated “model accuracy.” The long tail wasn’t hidden; it was routed to review or shown as an explicit unattributed line.

From project to operating model

I built a repeatable capability, not a one-off run

A senior program isn’t “get it working once.” It’s a mechanism that lands reliably every cycle and keeps earning trust. I converted the technical solution into a standing monthly operating model with validation gates and a managed backlog.

Artifact · Monthly operating cyclerefresh → consumption → backlog
Data refresh→ Pipeline execution→ Match-quality validation→ Spend reconciliation→ Pipeline reconciliation→ Exception review→ Tableau publication→ Leadership consumption→ Enhancement backlog
Operating modelcadence · gates · tools
DimensionHow I ran it
CadenceA monthly refresh + QA cycle — re-run the pipeline, validate match quality, reconcile spend and pipeline, publish the updated Tableau suite.
Validation gatesMatch-quality validation and spend/pipeline reconciliation had to pass before anything was published — no unreconciled number reached leadership.
Exception handlingBorderline matches worked through the review queue; unresolved records carried forward as the explicit unattributed line.
Backlog & roadmapRan the roadmap, refresh schedule, and enhancement backlog in Asana, and mapped it into the wider program plan in Microsoft Project so it stayed visible alongside the broader partner-marketing calendar.
Stakeholder commsTailored the read per audience and surfaced match confidence in the dashboard — credibility comes from showing the seams, not hiding them.

The takeaway

I converted a technical solution into a repeatable operating capability — the same trusted number, every month, with the gates and backlog that keep it trustworthy as the data and the questions evolve.

Outcomes

What changed — at four levels

I separate what was delivered from what the business could decide — and I’m deliberate about not inflating either.

1 · Program outputs — what was delivered

One unified model

Three enterprise systems — partner marketing, Salesforce, SAP — joined into a single MDF attribution model.

460K+ touches resolved

Free-text, multilingual partner-marketing touches resolved to canonical CRM accounts, with confidence and an honest unattributed line.

A standing capability

A monthly-refresh Tableau suite with reconciliation gates, an exception process, and a managed enhancement backlog.

2 · Technical outcomes — what changed in the data capability

Comparable, global data

Seven currencies to USD and years of free-text collapsed into a stable taxonomy — regions finally comparable like-for-like.

Defensible attribution

~82–85% of touches auto-resolved, on an explainable rules-based model where every credited dollar traces to a touch and account.

Visible uncertainty

Match confidence and an explicit unattributed slice built into the output — trust through transparency, not false precision.

3 · Business outcomes — what leadership could now decide

MDF ROI, finally

Spend tied to sourced and influenced pipeline — by partner, geo, activity, and solution — instead of disconnected spreadsheets.

Efficiency, not just totals

Cost-per-opportunity and cost-per-won-opportunity showed which partners and activities returned most per MDF dollar.

Promise vs. reality

A “solutions promoted vs. products actually purchased” view exposed where funded campaigns and real buying diverged.

4 · Operating outcomes — what became sustainable

Repeatable, not heroic

A monthly operating cycle anyone could rely on — not a number that was true once and stale after.

Trust that compounded

The same number meant the same thing in every region’s view, cycle after cycle — which is how distrust turned into reliance.

A managed backlog

Enhancements prioritized and delivered over time, so the capability improved instead of decaying.

What leadership saw — illustrative region view

The shape of the monthly MDF-ROI view partner leaders used to decide where the next dollar went — MDF invested against sourced pipeline, opportunities, and cost-per-opportunity, by partner. The numbers below are illustrative sample values to show the layout, not validated enterprise results.

MDF ROI & Sourced Pipeline · Region viewILLUSTRATIVE · SAMPLE VALUES
MDF invested$1.8M
Sourced pipeline$11.6M
Opportunities320
Cost / opportunity$5.6K
Pipeline : spend6.4×
Sourced pipeline by partner — top 6 (sample)
Partner A
Partner B
Partner C
Partner D
Partner E
Partner F

The honest version of the impact

I won’t attach an inflated ROI multiple to this. What’s true and defensible: partner leadership went from no trustworthy MDF attribution to a single monthly view — 460K+ touches resolved, three systems and seven currencies reconciled — that they used to judge partner and regional performance and steer where MDF gets invested. The sample dashboard figures above illustrate the layout, not a validated business result.

What this program taught me

Five lessons I carry into bigger programs

Technical deep dive · under the hood

For the engineering reader — how the core actually works

The program story stands on its own above. This section is for a technical reviewer who wants to confirm the depth is real. Expand what you want to inspect.

Identity-matching pipeline — step by step
StepWhat it doesWhy
CleanseStrip legal suffixes and punctuation; remove any word appearing in >0.5% of names.Removes noise (“Inc”, “GmbH”, “Technologies”) that would inflate false matches.
StandardizeConvert every country to a clean ISO-2 code; translate non-Latin names to Latin script.Lets an APAC touch and a US account name be compared on the same footing.
BlockOnly compare candidates within the same country (sorted-neighbourhood indexing).Cuts the comparison space by orders of magnitude — accuracy and compute cost.
ScoreSix similarity metrics — Jaro-Winkler, Damerau-Levenshtein, token-set and partial ratios over two cleansing passes — combined by harmonic mean.Harmonic mean punishes any single weak signal, so one flattering metric can’t carry a bad match.
ResolveAdd a composite city/country tie-break, rank, and keep the single best account per touch.One touch resolves to exactly one account — no double-counting downstream.
Harmonization & spend logic

Before anything could be matched, the raw claim data had to be made trustworthy. In BigQuery I parsed a delimited partner-level string into a clean geo → country → area hierarchy across the three theaters; wrote MDF-spend logic tied to each claim’s lifecycle status (in-flight claims use the requested amount, closed claims use the paid amount); and collapsed years of inconsistent free-text — activity types, business objectives, product and solution names — into a stable taxonomy. Unglamorous, but every downstream number depends on it.

Attribution model — touch → account → pipeline

The unit: one MDF-funded partner-marketing touch (event, webinar, telemarketing push, or campaign), carrying its funding claim and spend. The join: each touch is resolved to a canonical CRM account, then joined to that account’s pipeline and opportunities. The credit rule: a touch that reached an account is credited with the pipeline that account generated — sourced or influenced — in the window; because one touch resolves to exactly one account, nothing is double-counted. The efficiency lens: MDF investment ÷ opportunities from touched accounts gives cost-per-opportunity and cost-per-won-opportunity — the numbers leaders use to compare partners and activities. It’s deliberately a rules-based, account-based model, not a statistical one, so every credited dollar is explainable.

Confidence thresholds & exception handling
Automated match rateRoughly 82–85% resolved confidently through automated fuzzy matching and domain pairing on clean or semi-structured fields — using tokenization, edit-distance scoring, and normalization for legal suffixes, tax/VAT identifiers, and email domains.
Long tail (15–18%)Multi-byte characters in APAC claims, non-standard abbreviations, and records missing a domain or tax ID couldn’t resolve safely without risking false positives.
Human-in-the-loopRecords scoring 0.60–0.79 routed to an exception queue for regional partner operations to verify manually — borderline cases went to a person, not an algorithm guessing.
Below 0.60Tagged “Unattributed Partner Spend” rather than force-matched — shown as its own line so leadership sees the true size of what couldn’t be attributed.
Non-Latin handling & blocking

Non-Latin-script markets (Japan, Korea, Taiwan, Israel) were pulled and translated on a separate path rather than translating everything — translation is the most expensive step, so isolating the small share of records that actually need it kept the monthly refresh fast and cheap without losing coverage. Blocking by country before fuzzy comparison then cut the candidate space dramatically, which improved both accuracy (no cross-country false matches) and compute cost.

BigQueryPython · recordlinkage / fuzzywuzzySalesforceSAPTableauAsanaMicrosoft ProjectSQL