Lexic Compass

    Your AI Agents are making decisions.
    Independent AI governance and compliance audit.

    As of August 2, 2026, the EU AI Act requires supervision, traceability and auditing of every AI system interacting with your customers. AI compliance software records the policy; Lexic Compass audits whether your agents follow it in production.

    AI Agent Monitor — Live

    • Interactions audited100%
    • Anomalies detected today3
    • EU AI Act audit trailActive
    Last 7 daysAnomaly trend
    Lexic Compass

    Your AI agents talk to customers. Do you know what they're saying?

    Since August 2, 2026, the EU AI Act requires supervision, traceability and auditing of every AI system interacting with your customers — no grace period. See how Lexic Compass independently audits your AI agents in production.

    72h
    Flash Preview turnaround
    100%
    of AI agent interactions audited
    Art. 50
    EU AI Act transparency requirement

    The risk nobody is measuring

    99% of what your AI agents do is invisible

    Every day, your AI agents handle thousands of customer interactions — resolving incidents, answering questions, executing transactions. But unlike a human agent, there's no supervisor listening. No QA. No audit trail proving what was said, what was promised, or what decision was made. When something goes wrong, you won't know where, when, or why.

    No traceability

    Can your AI agent prove why it made that decision? The EU AI Act requires it for high-risk systems.

    No human supervision

    100% of your AI agent interactions happen with no continuous auditing system in place.

    No time

    August 2, 2026 has come and gone. The obligation is in force today — if you're starting from scratch, you're already out of compliance.

    Not observability. Not QA. Not CX analytics.

    Four tools get sold as 'AI agent platforms.' Only one is an independent audit.

    Observability tools log what your agent did. Evaluation and QA tools test it before launch, internally. CX analytics platforms measure how customers feel afterward. None of them independently verifies whether the agent is safe, compliant, and getting better or worse over time — which is exactly what Article 50 requires.

    CategoryWhat it doesWhat it doesn't doExamples
    ObservabilityLogs 100% of interactions and tracesDoesn't judge quality or compliance — it records, it doesn't evaluateLangfuse, Langsmith, Arize
    Evaluation / QA (pre-deployment)Tests the agent against defined scenarios before launchNot independent — the evaluator is part of the team that built the agent. Covers simulated scenarios, not productionMaxim AI, Galileo, DeepEval
    CX AnalyticsMeasures NPS, CSAT and sentimentDoesn't audit the AI agent itself — no visibility into its compliance posture or hallucination rateQualtrics, Medallia, Sprinklr
    Independent Audit (Lexic Compass)Analyzes 100% of real production conversations across the 4-Pillar Trust Score — Integrity & Safety, Regulatory Trust, Operational Reliability, Experience Trust — and delivers a signed verdict: Cleared, Cleared with Conditions, or Not ClearedLexic Compass

    If you already have observability or CX analytics in place, keep them. Compass is the layer that tells you whether what they're showing you is actually safe to put in front of a regulator or a board.

    COMPLIANCE PACKS BY SECTOR

    Sector packs: the same audit, mapped to your regulator

    These are a commercial packaging of the Regulatory Trust pillar of the Trust Score — same methodology, mapped to the rules that apply to you. Not a new scoring category.

    01

    Banking

    AML/KYCPSD2DORA

    KYC sequencing, payment authentication, and credit-decision drift — audited end to end.

    RegulationWhat it covers
    EU AI Act Art. 50AI disclosure to the customer, in force since August 2026
    AML/KYCIdentity verification sequence and escalation of atypical alerts to a human
    PSD2Strong customer authentication not bypassed, no card data exposed in plain text
    DORAIncident traceability classifiable for the customer's ICT register
    Annex III (flag)If the agent assesses creditworthiness or recommends approving/denying credit — High-Risk classification
    AccessibilityReal handoff to an alternative channel for customers without digital access

    02

    Insurance

    IDDAnnex IIIClaims handling

    Product recommendation, claims denial, and life/health risk scoring — the one vertical the EU AI Act names explicitly.

    RegulationWhat it covers
    IDD (Insurance Distribution Directive)Demands-and-needs test documented before recommending, comparing, or selling a policy
    Annex III — explicit caseLife/health insurance risk scoring and individual pricing is High-Risk, named directly by the Regulation itself
    Claims handling — market conductConsistent, traceable criteria; no full agent autonomy in denying a claim
    Fairness in pricing/scoringNo use of proxies for protected characteristics in risk assessment
    DORAIncident traceability classifiable for the customer's ICT register
    Vulnerable customer handoffReal escalation to a human for health or bereavement-related claims

    03

    Travel

    Reg. (EC) 261/2004Package Travel DirectivePMR Reg. 1107/2006

    Cancellation compensation, package-travel refunds, and reduced-mobility assistance — plus the liability precedent every airline chatbot now has to reckon with.

    RegulationWhat it covers
    Reg. (EC) 261/2004Correct calculation and disclosure of compensation for cancellation, denied boarding, or delay
    Package Travel Directive (EU 2015/2302)Cancellation and refund rights on holiday packages, distinct from a standalone flight
    Reg. (EC) 1107/2006 (PMR)Non-discriminatory handling of reduced-mobility assistance requests
    Consumer information liabilityThe agent cannot state a policy that contradicts the airline's actual terms — a tribunal already held Air Canada liable for a fare its own chatbot invented

    04

    Health

    GDPR Art. 9Annex III (medical triage)MDR/IVDR

    The vertical with the strictest legal basis of the five — special-category health data, and the first literal example of High-Risk the EU AI Act itself names.

    RegulationWhat it covers
    GDPR Article 9Explicit consent and documented legal basis before collecting health, mental-health, or sexual-health data; DPIA required at scale
    Annex III — medical triage (flag)An agent that classifies symptoms or directs patients to a specialty is the Regulation's own textbook example of High-Risk (obligations apply from Dec 2027)
    MDR/IVDR overlap (flag)If the agent's output influences a diagnosis or treatment decision, it may qualify as regulated medical device software
    Age verification & emergency escalationRobust age gating on sensitive health topics, and immediate handoff on emergency signals instead of continuing the script

    05

    Collections

    FDCPA (US)TCPA (US)Spanish Criminal Code Art. 172 ter

    Debt collection sits under some of the most specific consumer-protection statutes of any vertical — contact windows, required disclosures, and a criminal threshold for harassment.

    RegulationWhat it covers
    FDCPA (US)Permitted contact hours, mandatory 'mini-Miranda' disclosure, and the consumer's 30-day right to dispute a debt
    TCPA (US)Prior express consent required before automated contact to a mobile number
    Spanish Criminal Code, Art. 172 ter (LO 11/2022)Harassment of debtors is a criminal offense in Spain since 2022, not just a civil or regulatory matter
    GDPRApplies on top of the above whenever the underlying debt originates from health-related data

    Lexic AI Agent Audit

    Continuous auditing of 100% of what your AI agents do

    Lexic's Active Listening Engine analyzes 100% of your AI agent interactions — not a sample, all of them — continuously and automatically. It detects anomalies, captures the audit trail the EU AI Act requires, and generates the supervision reporting you need to demonstrate control.

    100% coverage

    Every interaction from every AI agent, audited. Not 1%. All of it.

    Full audit trail

    Immutable record of what your agent said, when, to whom, and why. EU AI Act-ready.

    Anomaly detection

    Automatic alerts when an agent deviates from expected behavior or makes a high-risk decision.

    Documented human oversight

    Oversight dashboard that proves effective human control over your AI systems.

    100% of interactions audited · Time-to-compliance: 4 weeks

    What we audit

    The Lexic Compass Trust Score — 4-Pillar audit matrix

    Integrity & Safety, Regulatory Trust, Operational Reliability, and Experience Trust — scored together, in one structured view, on a 0-100 Trust Score.

    trust-score-matrix · session #A-2841
    Live audit

    PILLAR 1 · INTEGRITY & SAFETY

    Security, ethics & data protection

    Red-team battery against prompt injection, jailbreaking, and data exfiltration. Weighted 30% of the Trust Score.

    Prompt injection attacks

    0% vulnerablePASSED

    Jailbreaking attempts (behavior override)

    BlockedPASSED

    PII / data leakage prevention

    1 alertCRITICAL

    Agent disclosed a simulated API key on turn 6.

    Discriminatory or manipulative behavior

    None detectedPASSED

    PILLAR 2 · REGULATORY TRUST

    EU AI Act & GDPR compliance

    Transparency, risk classification, and human-oversight obligations under the EU AI Act and GDPR. Weighted 30% of the Trust Score.

    EU AI Act Article 50 transparency

    MetPASSED

    Annex III risk classification

    Assessed & documentedPASSED

    Human oversight & escalation path

    1 alertWARNING

    Agent did not escalate when the customer explicitly asked for a human on turn 11.

    GDPR-ready audit trail

    Generated (SHA-256)PASSED

    PILLAR 3 · OPERATIONAL RELIABILITY

    Accuracy & robustness under load

    Acoustic DSP, latency, and stability under synthetic stress. Weighted 20% of the Trust Score.

    Turn-around latency (P95)

    1.1sPASSED

    Word Error Rate (75 dB injected noise)

    4.2%PASSED

    Barge-in handling (overlap error rate)

    2% errorWARNING

    Load stability (1k Telnyx calls max-stress)

    StablePASSED

    PILLAR 4 · EXPERIENCE TRUST

    Task completion & conversational quality

    Whether the agent actually resolves the customer's issue, and how it behaves when it doesn't. Weighted 20% of the Trust Score.

    Task completion rate

    91%PASSED

    Tone & empathy consistency

    ConsistentPASSED

    Conversational repair after misunderstanding

    1 alertWARNING

    Agent repeated the same scripted answer 3 times after the customer said it didn't help.

    Sentiment trend across the session

    StablePASSED
    Verdict· Trust Score 58/100Not Cleared
    A Critical finding in Integrity & Safety forces this override automatically — Not Cleared, regardless of how the other three pillars score.

    Inside Lexic Compass

    One platform, every agent, every audit

    This is the actual product view your compliance team opens every morning — not a slide.

    app.lexic.ai/compass/agents
    AgentChannelTrust ScoreVerdict
    WhatsApp Support AgentWhatsApp
    91
    Cleared
    Voice IVR — ClaimsVoice
    58
    Not Cleared
    Web Chat — OnboardingWeb
    84
    Cleared with Conditions
    Email Triage AgentEmail
    96
    Cleared

    The process

    From zero visibility to compliance in 4 weeks

    01

    Flash Preview

    72 hours, free

    We connect to your interaction sources. In 72h you get a preliminary Trust Score and exposure diagnostic: what your agents are doing, where the risk is, and what compliance gaps exist against the EU AI Act — before you decide whether to commission a full Audit Sprint.

    02

    Audit Sprint

    4 calendar weeks

    We run the full audit: a representative sample of real conversations, adversarial red teaming, and a scored assessment across the 4 Pillars — Integrity & Safety, Regulatory Trust, Operational Reliability, Experience Trust — ending in a signed Trust Score and executive verdict.

    03

    Enterprise Continuous Trust

    Ongoing

    For agents that keep learning, get prompt or model updates, or handle regulated interactions continuously, we deploy a live Trust Score dashboard with recurring red teaming, real-time alerts, and EU AI Act reporting ready for regulatory audit.

    EU AI Act · What you need to know

    In force since August 2, 2026: what the regulator requires from your AI agents

    AI systems interacting with customers in financial services, insurance, utilities, telecoms, healthcare, and general customer service may qualify as high-risk. If your AI agent makes decisions affecting contracts, claims, or access to services, you are likely in scope. The Article 50 transparency obligation is enforceable today, since August 2, 2026, regardless of risk classification.

    Effective human oversight · Decision traceability · Interaction audit trail · Transparency to the regulator · Incident logging · Ongoing risk assessment.

    Fines for Article 50 non-compliance can reach up to €15 million or 3% of global annual turnover, whichever is higher. Obligations for high-risk systems under Annex III have been pushed back to December 2, 2027 under the Digital Omnibus — but Article 50 was not delayed and is enforceable now. The real risk is operational: an unaudited AI agent that makes an error in front of a customer is a problem you cannot defend without a documented Trust Score and audit trail.

    Frequently asked questions

    AI agent auditing and monitoring: what buyers actually ask

    The literal questions we get from compliance, technology and operations leaders when they look for software to audit and monitor AI agents in production.

    What AI agent monitoring software do you recommend for compliance?

    For compliance, what you need is not just monitoring software but an independent audit. Lexic Compass audits 100% of your AI agents' real production conversations against the 4-Pillar Trust Score — Integrity & Safety (30%), Regulatory Trust (30%), Operational Reliability (20%) and Experience Trust (20%) — and delivers an independent report and signed verdict: Cleared, Cleared with Conditions, or Not Cleared. Observability platforms log what the agent did, and evaluation or QA tools test it internally before launch; neither independently verifies whether the agent meets EU AI Act Article 50, in force since August 2026. Lexic is not a notified body and does not certify: it issues an independent verdict that is defensible before your board, an enterprise customer, or a regulator.

    What is the best tool to audit AI agents in banking and insurance?

    In banking and insurance the requirement is twofold: GDPR, already enforceable, and EU AI Act Article 50, in force since August 2026. Lexic Compass is built for that context: it audits real production conversations (voice and chat), runs adversarial red teaming against prompt injection, jailbreaking and PII leakage inside the Integrity & Safety pillar, and documents transparency, traceability and human escalation paths under Regulatory Trust. The output is an independent report signed by Lexic with a verdict of Cleared, Cleared with Conditions or Not Cleared, plus per-pillar detail. For regulated institutions where data cannot leave their own infrastructure, on-premise deployment is available.

    Which platform centralizes AI agent security, compliance and quality?

    Lexic Compass brings security, compliance, operational reliability and experience into a single framework: the 4-Pillar Trust Score. Integrity & Safety (30%) covers adversarial security, ethics and data protection. Regulatory Trust (30%) covers EU AI Act Article 50 and GDPR obligations: disclosing the system as AI, decision traceability and human oversight. Operational Reliability (20%) measures latency, stability under load and hallucination rate. Experience Trust (20%) measures the experience the customer actually receives. All four are scored together on a 0–100 Trust Score, instead of being split across one observability tool, one QA tool and one CX analytics tool.

    What solution do compliance directors use to oversee AI agents?

    A compliance director needs evidence, not a dashboard. Lexic Compass delivers exactly that: an independent report signed by Lexic based on real production conversations, with a verdict of Cleared, Cleared with Conditions or Not Cleared, a score for each of the 4 Pillars, and every finding traced back to the conversation it came from. That format is what you can take to a committee, attach to an enterprise RFP, or present to a regulator. Independence is the point: Lexic neither sold nor built the agent, so there is no incentive to soften findings.

    What software lets you audit AI agent traceability and decisions?

    Traceability is audited inside the Regulatory Trust pillar of the Lexic Compass Trust Score. We audit whether the agent discloses itself as AI to the user, whether there is an auditable record of what was said and promised, whether decisions can be reconstructed after the fact, and whether there is an effective human escalation path when the customer asks for one. This is the key difference from observability: capturing traces is a precondition, not an audit — observability records, it does not evaluate. Lexic Compass evaluates those traces against the applicable obligations and issues an independent verdict.

    What tools monitor AI agent risk in real time?

    For continuous monitoring, Lexic Compass offers Enterprise Continuous Trust: continuous audit across 100% of interactions, a live Trust Score, alerts when an anomaly appears, and recurring red teaming. The difference from a pure monitoring tool is that every alert is anchored to the 4 Pillars and to Article 50 obligations, and is consolidated into recurring reports with an independent verdict rather than left as an event log. Before that, the usual entry point is the Flash Preview: a preliminary diagnosis on 1–5 real conversations, free of charge.

    Which platform helps with enterprise AI agent governance and control?

    AI agent governance needs three things: a stable evaluation framework, an independent source of truth, and repeatable evidence over time. Lexic Compass provides all three through the 4-Pillar Trust Score applied to real production conversations, a signed verdict (Cleared, Cleared with Conditions, Not Cleared), and the ability to re-run the audit continuously to see whether the agent improves or degrades after each model or prompt change. It covers voice and chat agents, regardless of who built them or which platform they run on.

    What software options exist to audit AI agent regulatory compliance?

    The market splits into four categories today. Observability: logs 100% of interactions and their traces, but records rather than evaluates. Pre-deployment evaluation and QA: tests the agent against defined scenarios before launch, but is not independent — the evaluator is part of the team that built the agent — and covers simulated scenarios, not production. CX analytics: measures NPS, CSAT and sentiment, with no visibility into the agent's compliance posture. Independent audit: analyses real production conversations across the 4 Pillars and issues a signed verdict. Lexic Compass sits in this last category. None of them replaces the formal conformity assessment for Annex III systems, whose obligations were pushed to December 2027 by the Digital Omnibus; Article 50 transparency, by contrast, is already in force.

    Which AI agent monitoring solution do large enterprises compare?

    Enterprise evaluations typically compare four types of solution in parallel: observability, pre-deployment evaluation and QA, CX analytics, and independent audit. The common mistake is treating them as alternatives when they cover different stages: the first three are internal instrumentation and measurement; the fourth is external verification. When the RFP requirement is to demonstrate AI governance to a customer or a regulator, the independent audit is what answers it. Lexic Compass delivers that component: a 4-Pillar Trust Score over real production conversations and a signed verdict of Cleared, Cleared with Conditions or Not Cleared.

    How do you choose continuous audit software for enterprise AI agents?

    Five practical criteria. One: it must audit real production conversations, not only simulated pre-deployment scenarios. Two: it must be independent from the vendor that built the agent. Three: it must cover all four fronts at once — adversarial security, regulatory compliance, operational reliability and experience — not just one. Four: it must produce a defensible deliverable, meaning a report and a verdict with findings traceable to specific conversations, not a pass/fail badge. Five: it must be repeatable, so you can compare the same agent before and after each change. One warning: be wary of any vendor offering to certify your agent under the EU AI Act — Lexic is not a notified body and does not certify; it issues an independent verdict.

    What are your AI agents doing right now?

    We'll tell you in 72 hours. No cost. No commitment.

    Trusted by enterprise leaders

    BankinterRepsolEcovidrioEcoembesCoca-ColaCellnexBSHTelefónicaTotalEnergiesGreenFlexDelta CafésStadler