Papanikolau and Papanikolaou represent the same surname, differing by a single letter due to alternate Greek transliterations. Christian and Christina share identical letters, yet represent entirely different given names. A compliance analyst grasps this context instantly. An edit-distance algorithm merely counts character shifts, unable to tell a spelling variant from a completely different identity.
Sama™ is AMLYZE’s answer to that gap: a machine learning model, built in-house and running inside its screening engine since September 2026, that judges word by word how closely two names refer to the same person, and shows its working on every match.
Dominykas Stankevičius, Head of Data Science and Analytics at AMLYZE, explains why it exists, how it works, and why the team spent far more time on the data than on the model.
Dominykas, let’s go straight to it. What is Sama?
A machine learning model we trained ourselves, doing one very narrow job: you give it two words, it tells you how likely it is that they’re the same name. That’s the whole product. It produces a score, and the analyst takes it from there. Built in-house, our own data, our own pipeline, all of it ours.
And what was broken that made you build it?
Name screening is a permanent trade-off. Loosen your matching, and you drown in false positives. Tighten it, and you start missing real hits, which is the one thing you cannot do in this job. So most teams pick the safe side and accept the noise, and your best people spend the day clearing alerts that were never going to be real hits. Our engine already runs below a 1.00% false positive rate in production, with more than a hundred sources behind it and over forty tunable parameters. But below one percent of a very large number is still a lot of names on somebody’s screen every day, and the ones that survive to that screen are the hard ones.
Hard how? Give us something real.
Very often these are names that have travelled between alphabets. Take a Belarusian patronymic: one list writes it Uladzimiravitj, another writes the same patronymic Vladimirovich. Neither is wrong, they just came through different transliteration conventions, and now you have a handful of letters of difference on what is one person. Or a Georgian surname like Msjvidobadze, where the convention alone changes several letters. An analyst looks at that for two seconds and usually knows. A classic edit-distance-based algorithm has no intuition, it counts how many characters you’d have to change, with no idea whether that change is a normal transliteration or a completely different person.
People hear “AI screening” and assume you threw out the old engine.
Worth being precise here, because it’s the thing I most want people to understand. The engine keeps doing everything it did before: it finds the candidates, handles the structure of the name, applies the clearance rules and the penalties. Sama makes one step inside it better, the word level. We compare two names word against word rather than as one long string, and Sama scores how close two words are to being the same name word. And it works exactly where that judgement is hardest. The engine scores a word pair somewhere between zero and one, and Sama is asked only about the ones that land close to a match without quite being one. Below that they’re clearly different words, above that clearly the same. It’s the narrow strip just below a clean match where an extra signal earns its place, and in practice only a small share of hits ever land there.
Explain the model itself like we’re not technical. And where does it learn all this?
Think about what you do when two names sit side by side. Your eye finds the differences on its own, and then something you’ve learned decides whether they matter: is the root the same, is one just longer, is this the kind of ending that changes when a name crosses a language. Nobody ever taught you that as a rule. You built it up from seeing a lot of names, and from knowing something about where they come from. Sama is a neural network, and a neural network is simply a way of letting software build that same judgement out of examples instead of being handed the rules.
So take a real one: is Mohammed the same name as Muhamad? The smallest piece of the network, a neuron, is just a small sum. It takes a few things it can measure, do the first letters match, how many letters do the two words share, how different are the lengths, multiplies each by a weight that says how much that clue counts, and adds them up into one number. Sama is that same idea repeated 1,024 times, with 141,904 weights, arranged so it reads a word letter after letter rather than all at once. Nobody sets those weights by hand. The model finds them by working through millions of examples, which is the same way you got your instinct, only faster and a great deal more literal.
The training set comprises about 3.8 million name pairs: 2.9 million from our own candidate data, 640,000 watchlist aliases, 258,000 independently confirmed aliases, and 13,000 synthetic pairs covering rare but real manipulations. We quarantine untrusted word pairs rather than feed them into the model. Training runs daily – while one model serves traffic, another is being built -but that’s the cheap part. The data is the expensive part. Always.
The data is the expensive part. Always.
Is it trained on customer data?
No, not unless a client specifically agrees to it. The model is trained on watchlist material, the aliases published by the lists themselves, and variations we generate internally. Even when a client does agree, we never take their raw customer records. Instead, we use their feedback: whether a match was flagged as a false positive or a true positive. That verdict tells us something no static dataset can. We take that signal in an anonymized form – stripping away entity details to keep only distinct words from the name – and use it to train the model on the exact correction the analyst just made. If a client prefers not to participate, nothing changes for them: they get the same model as everyone else.
Is it finished, or do you keep training it?
Every night. A candidate model comes out, gets evaluated automatically against whatever is live, and the verdict is improved or not improved. Most nights it’s not improved, and that’s the system working. The testing runs on data the model has never seen, around half a million pairs held back from training for exactly that purpose. The numbers we watch hardest come from the slice of that set inside the band where Sama actually runs, 11,444 pairs: there it catches 97.3% of the pairs that genuinely are the same name spelled differently, and when it calls a pair the same name it’s right 95 times out of 100. We keep confidence intervals around those, and tolerance bounds so we don’t promote a model every time the numbers wobble in the third decimal.
And what would that mean for the alerts a team actually sees?
We started Sama on a deliberately narrow scope, because that’s how you keep the result under control. It’s consulted only on the hits where the conventional engine is least certain, and everywhere else the engine’s answer stands untouched. Then we replayed it over real screening hits to see what happens inside that scope. Of the hits Sama reviews, about seven in ten would shrink or disappear: sixty-nine out of a hundred stop being a match, two keep a smaller set of matched records, and twenty-nine look exactly as they do today. And the scope isn’t fixed. With each iteration the range widens, and more of the false positives a team clears by hand come within reach.
Here’s the pushback. Somebody’s going to say this is fuzzy matching with an AI sticker on it.
It’s a fair question, and one we put to ourselves. Fuzzy matching runs a fixed formula, count the edits, output a number. It has no concept of whether the change it’s looking at is a plausible transliteration or a completely different person. Sama learned that distinction from data, and you can see it: the two give different answers precisely on the cases where it matters, and Sama’s is the one closer to what an analyst would say.
It worked out on its own that vowels are where spelling drifts and consonants are where identity lives.
Anything it learned that you didn’t teach it?
This is one of my favorite insights. We never gave the model a single phonetic rule, yet it behaves as if it knows them. We test this by taking a real name, altering a single letter, and measuring how much the similarity score drops. Change a vowel and the score barely moves – dropping only about 3 percentage points – while swapping an i for a y leaves an almost perfect match. Change a consonant, however, and the score collapses by at least half. It worked out on its own that vowels are where spelling drifts and consonants are where identity lives, which is roughly how a human reads an unfamiliar name.
What else did those tests tell you?
Two things I like. The first letter is sacred: replacing it gives the lowest score of every test we run, and in seven of the eight tests where we checked, the same edit hurts more at the start of a word than anywhere else. And longer words absorb damage, so short names are the fragile ones, which matches the intuition that you have more to go on with Uladzimiravitj than with Li. What barely matters is language. Only six of nineteen tests showed any detectable difference between language groups.
Where does it make the biggest difference?
The biggest impact is single consonant change handled differently in different scenarios. Genuinely distinct names may vary per single consonant, which for analysts is an obvious mismatch, but for classic algorithms it’s no difference from translation from one alphabet to a language using latin alphabet. Sama has learnt to differentiate these two diverse situations. Then there are names where the variation could be a genuine alias or two unrelated people, and the name alone can’t settle it. Sama gives a middle score there, which is exactly what you want, because that’s where an analyst’s judgement belongs. And with long names where several things changed at once, it’s one signal among several, alongside the date of birth, the country and the identifiers.
Same name or different person? How the conventional engine and Sama score six real name pairs
| Watchlist name | Screened name | Same name? | What changed | Engine score | Sama score |
|---|---|---|---|---|---|
| Papanikolaou | Papanikolau | Yes | Greek "ou" written as "u" | 0.92 | 0.99 |
| Guluashvili | Guluasjvili | Yes | Georgian "sh" written as "sj" | 0.91 | 0.99 |
| Aliaksandr | Alyaksandr | Yes | i / y transliteration | 0.9 | 1 |
| Mshvidobadze | Mshvitobadze | Yes | d / t, similar sounds | 0.92 | 0.86 |
| Shatkovskaya | Shatrovskaya | No | one consonant changed | 0.92 | 0.11 |
| Martinas | Martina | No | male vs female form | 0.88 | 0.39 |
Compliance officers are allergic to black boxes, and for good reason.
We set one condition before any model went into the platform: a compliance officer has to be able to see why a match scored the way it did and explain it to a regulator. So on every match you get the full breakdown, which input word was compared against which watchlist word, what Sama scored that pair, what the conventional engine scored the same pair, and every rule applied on top. Both numbers, side by side, are labelled, so you can reconstruct the arithmetic yourself. I’d be precise about what that means: you can’t read the inside of a neural network the way you read a rule, and I wouldn’t claim otherwise. What you can do is see every number that went into the result, test the model’s behaviour and show your working. And the decision stays with people: Sama produces a score, the engine produces a match, the analyst reviews the alert and decides.
And you can run it without letting it touch your scoring at all.
That’s the default. In advisory mode Sama’s score just sits next to the existing one, as information, and your scoring and alert generation are completely unchanged. It costs you nothing to watch it for a few weeks on your own data. Only when you’ve seen enough do you move it into primary, where its result becomes the effective name score. Either way it’s opt-in, and either way the analyst still reviews every alert.
I’d rather lose a week to that question than ship something I can’t explain six months later.
What was the hardest part of building it?
The one I remember is a model that came back better than it had any right to be. The numbers jumped, everyone was pleased for about an hour, and then we started asking why. A result that good usually means the model has memorised something it shouldn’t have, or that something from the test set has leaked into the training set. So instead of promoting it we spent days taking it apart, changing how we evaluate and when we stop training. It turned out to be fine. But I’d rather lose a week to that question than ship something I can’t explain six months later, and that’s the part of this work nobody puts in a demo.
Last one. What would you say to someone who’s cautious about AI in AML?
Turn it on in advisory mode and watch it. Your scoring stays exactly as it is, and in a few weeks you’ll have your own evidence from your own data. That’s a much better basis for a decision than anything I can tell you in an interview. And whatever we build after Sama, the standard stays the same: a compliance officer has to be able to explain the result to a regulator.
—
Test your own screening
Dominykas’s advice holds beyond Sama: the best basis for a decision is how your system performs on your own data.
Our Sanctions Screening Performance Test gives you that view. It runs controlled and manipulated test identities through your existing screening, including transliterations and spelling variants like the ones in this interview. The result shows what your system catches, what it misses, and where the noise comes from. No production data or system changes are needed.
Readers of this interview get 20% off the test. Just mention code SAMA26 when you request a scoping call. The offer is valid until 1 December 2026.





