One Search Query to Find Any User Across Every Directory
Most identity outages are not authentication failures. They start when your team cannot answer a simpler question fast enough: "Where does this user exist?" This guide shows how to build one search query that finds any user across Active Directory, Entra ID, LDAP, Okta, Google Workspace, and HR systems without forcing a full identity rewrite.
Nesqual Tech AI
A single missed user record can turn a routine offboarding into a breach. In a 40,000-user enterprise, even a 0.1% directory mismatch means 40 identities your team may not see during an incident, audit, or access review.
The fix is not another admin portal. It is a unified user search query that resolves one person across every directory you run: Active Directory, Entra ID, Okta, LDAP, Google Workspace, and the HR system that started the identity lifecycle in the first place.
This post shows how to design that query, normalize identity data, and make the result fast enough for service desks and reliable enough for security operations.
Why one user search query beats six admin consoles
When an engineer leaves, your team does not care which directory is authoritative for displayName. You care whether the user still exists anywhere, which accounts are active, and what to disable first.
That is why the best enterprise pattern in 2026 is not "pick one directory." It is search once, resolve everywhere.
The failure scenario most teams recognize
Consider a real-world setup:
- Active Directory for legacy Windows workloads
- Microsoft Entra ID for Microsoft 365 and conditional access
- Okta for SaaS federation
- Google Workspace for a subsidiary
- Workday as the HR source
- Two LDAP directories still attached to old customer portals
A service desk agent searches j.smith in Entra ID and finds nothing. Security closes the ticket. Twelve hours later, the same user is still active in a legacy LDAP-backed VPN portal because the username there is jsmith_ext and the email is an old alias.
The problem was not provisioning. The problem was search fragmentation.
What the unified query must return
Your one search query should answer five questions in under two seconds:
- Does this person exist in any connected directory?
- Which identifiers match: username, email, employee ID, external ID, phone, alias?
- Which account is authoritative and which are linked shadows?
- What is the current state: active, suspended, disabled, deleted, pending?
- What systems still hold an account for this user?
For most enterprise support and security workflows, that is enough to cut mean time to identify by 60-80%.
Build an identity resolution model before you write the query
If you search raw directories directly, you will get noisy results, duplicate records, and false negatives. The query works only when you normalize identity first.
Start with a canonical user document
Create a canonical schema that every connector maps into. Keep it boring and explicit.
{
"personId": "8c3a8d7e-2f6b-4c4a-a1f0-9f7d8d2f4a11",
"authoritativeSource": "workday",
"employeeId": "104882",
"usernames": ["jsmith", "j.smith", "jsmith_ext"],
"emails": ["jane.smith@corp.com", "j.smith@subsidiary.com"],
"displayName": "Jane Smith",
"givenName": "Jane",
"familyName": "Smith",
"phones": ["+1-415-555-0148"],
"accounts": [
{"system": "active-directory", "accountId": "CN=Jane Smith,OU=Users,DC=corp,DC=local", "status": "disabled"},
{"system": "entra-id", "accountId": "0f2d...", "status": "active"},
{"system": "okta", "accountId": "00u8...", "status": "suspended"}
],
"aliases": ["Jane A Smith"],
"lastUpdated": "2026-09-18T10:42:11Z"
}
This document is what your unified search query hits. Not the source systems directly.
Match on stable identifiers first, fuzzy attributes second
The ranking order matters. If you treat displayName and employeeId as equals, you will create dangerous collisions.
A practical weighting model:
- Employee ID or HR person number: weight 100
- Immutable external ID from IdP: weight 95
- Primary email exact match: weight 90
- Username exact match: weight 80
- Proxy email or alias exact match: weight 70
- Phone exact match after normalization: weight 60
- Display name fuzzy match: weight 30
This lets 104882 beat Jane Smith, and jane.smith@corp.com beat jsmith when both exist.
Use a confidence score, not a binary match
A confidence score gives your service desk and SOC a way to trust the result. For example:
95-100: exact identity match, safe for automated workflows80-94: likely same person, show linked accounts and warnings60-79: possible match, require analyst confirmation<60: return as related result, not primary hit
This is the difference between a useful search tool and a liability.
The one search query pattern that works in production
The best pattern is a federated ingest, centralized search model. Pull data from directories on a schedule or event stream, normalize into a search index, then query that index with one endpoint.
Reference architecture
[Workday] ----\
[AD/LDAP] -----\
[Entra ID] ------> [Connector Layer] -> [Normalization + Identity Graph] -> [Search Index] -> [API/UI]
[Okta] --------/
[Google WS] ---/
[Legacy Apps] -/
Search flow:
User input -> query parser -> exact match pass -> alias expansion -> fuzzy pass -> confidence ranking -> result with linked accounts
This architecture avoids live-query bottlenecks. It also gives you a place to deduplicate and enrich records.
A practical query contract
Your API should accept one input and search across all mapped identifiers.
POST /v1/identity/search
Content-Type: application/json
{
"query": "j.smith@corp.com",
"tenant": "acme-prod",
"includeAccounts": true,
"maxResults": 10,
"fuzzy": true
}
And return ranked, explainable results:
{
"query": "j.smith@corp.com",
"results": [
{
"personId": "8c3a8d7e-2f6b-4c4a-a1f0-9f7d8d2f4a11",
"displayName": "Jane Smith",
"confidence": 97,
"matchedOn": ["emails.primary.exact", "accounts.okta.profile.login.exact"],
"accounts": [
{"system": "entra-id", "status": "active"},
{"system": "active-directory", "status": "disabled"},
{"system": "vpn-ldap", "status": "active"}
]
}
]
}
The matchedOn field matters. Engineers trust search results more when they can see why the engine made the match.
Query behavior that reduces false negatives
A good unified user search should perform these steps in order:
- Normalize input: lowercase, trim, Unicode normalize, strip punctuation where safe
- Detect query type: email, username, employee ID, phone, free text
- Run exact match against high-confidence fields
- Expand aliases and historical identifiers
- Run fuzzy match only if exact match confidence is low
- Return linked accounts grouped by person
That sequence keeps latency low and avoids burying exact hits under fuzzy noise.
Performance targets and benchmarks you should expect in 2026
If this search is for support, security, and IAM operations, speed is not a nice-to-have. Slow search drives people back to manual console hopping.
Realistic latency targets
For a centralized search index on OpenSearch 3.x, Elasticsearch 9.x, or a tuned PostgreSQL 17 + pg_trgm setup, these are practical targets for a 500,000-person enterprise corpus with 3-8 linked accounts per person:
- P50 exact search latency: 80-150 ms
- P95 exact search latency: 250-450 ms
- P95 fuzzy search latency: 500-900 ms
- Freshness SLA from source event to searchable record: 30-120 seconds
- Index storage footprint: 8-20 GB per 1 million normalized users, depending on retained aliases and audit metadata
If your search takes 3-5 seconds, users will stop trusting it for incident response.
Benchmark example: mixed-directory enterprise
In one representative architecture review, a 220,000-user organization indexed:
- 220,000 HR identities
- 248,000 Entra ID accounts
- 191,000 AD accounts
- 67,000 Okta users
- 34,000 Google Workspace users
- 420,000 historical aliases
With exact-first ranking and alias expansion, the team measured:
- 92% of service desk searches resolved on the first result
- 71% reduction in average lookup time, from 3.8 minutes to 66 seconds
- 84% reduction in missed secondary accounts during offboarding checks
- 0.6% false-positive rate on fuzzy-only searches after adding employee ID precedence
Those are the numbers executives care about because they map to ticket volume, audit findings, and risk.
Implementation choices: index, connectors, and ranking logic
You do not need exotic tooling. You need consistent connectors, a clean schema, and ranking logic your team can explain.
Connector strategy
Use the native APIs where available and LDAP only where necessary.
- Entra ID: Microsoft Graph delta queries
- Okta: Users API with event hooks for near-real-time updates
- Google Workspace: Admin SDK Directory API
- Active Directory and LDAP: incremental sync with
uSNChangedor equivalent change tracking - HR source: event feed or scheduled export with immutable person IDs
If a source cannot emit events, poll it. But record source freshness so operators know whether a result is stale.
Example normalization pipeline
import re
import unicodedata
def normalize_email(value: str) -> str:
return unicodedata.normalize("NFKC", value.strip().lower())
def normalize_username(value: str) -> str:
value = unicodedata.normalize("NFKC", value.strip().lower())
return re.sub(r"[^a-z0-9._-]", "", value)
def build_search_keys(user):
keys = set()
for email in user.get("emails", []):
keys.add(("email", normalize_email(email), 90))
for username in user.get("usernames", []):
keys.add(("username", normalize_username(username), 80))
if user.get("employeeId"):
keys.add(("employeeId", user["employeeId"], 100))
for alias in user.get("aliases", []):
keys.add(("alias", unicodedata.normalize("NFKC", alias.strip().lower()), 30))
return sorted(keys, key=lambda x: x[2], reverse=True)
This is simple by design. Most search failures come from inconsistent normalization, not weak infrastructure.
Ranking logic you can defend in an audit
Document your match precedence and keep it versioned. If the IAM team changes ranking, security and compliance should know.
A common pattern:
- Exact identifier match wins over all fuzzy matches
- Authoritative-source identifiers outrank downstream copies
- Active accounts rank above disabled accounts when confidence is equal
- Recent updates increase tie-break priority
- Multiple exact matches across systems increase confidence
That makes search behavior predictable during access reviews and forensic work.
Common Pitfalls
The hard part is not building search. It is avoiding the shortcuts that make it unreliable.
Treating email as immutable
It is not. Mergers, domain changes, and legal name changes break email-based identity matching.
Avoid it: use HR person ID or immutable IdP object ID as the root link key wherever possible.
Querying live directories at request time
This looks simple until one LDAP server slows down and your whole search path stalls.
Avoid it: index centrally and set freshness SLAs. For critical systems, add event-driven updates and a fallback cache.
Using fuzzy matching too early
If Jane Smith returns five people before exact email matches, your service desk will pick the wrong record under pressure.
Avoid it: exact-first, fuzzy-second, with confidence thresholds.
Ignoring historical aliases
Offboarding and incident response often depend on old usernames, old emails, and contractor IDs.
Avoid it: retain historical identifiers for at least 12-24 months, subject to your retention policy.
No explanation of why a result matched
A black-box result slows analysts down because they need to verify it manually.
Avoid it: return matchedOn, source freshness, and linked-account evidence in the response.
Forgetting authorization boundaries
A universal search API can leak more identity data than intended.
Avoid it: enforce field-level authorization. A help desk agent may need account status, while HR data should stay masked.
Key Takeaways
- Build a unified user search query on top of a normalized identity index, not by querying each directory live.
- Use a canonical person record with linked accounts, historical aliases, and an authoritative source field.
- Rank exact matches above fuzzy ones, and score stable identifiers like employee ID higher than names or aliases.
- Target sub-second P95 exact search latency and under two minutes freshness from source update to searchable record.
- Return explainable results with confidence scores,
matchedOnevidence, and linked account states. - This week, audit your current search paths: list every identity source, identify immutable IDs, and define one API contract your service desk and SOC can share.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Written by
Nesqual Tech AI
Nesqual Tech
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI