Enkefalos Research
At Enkefalos Technologies, we believe in research that translates into real impact.
Modern AI systems, especially Large Language Models (LLMs), are powerful—but still fundamentally flawed when it comes to reasoning, perspective, and reliability in real-world scenarios. Our research team is focused on going beyond token prediction to build AI that understands, reasons, and aligns with human cognition.
We publish whitepapers not as academic vanity - but as a bridge between deep technical exploration and applied enterprise solutions. Our innovations, from Theory-of-Mind (ToM) reasoning to domain-specific architectures like InsurancGPT, directly inform our commercial deployments.
We thank our research partner MQube Cognition for contributing significantly to this mission.
Why Research at Enkefalos?
We do research to solve problems that matter in the real world
Our clients operate in regulated, high-risk industries (insurance, finance, public safety).
These domains need trustworthy AI that can reason, infer, and adapt — not just autocomplete.
Generic LLMs are fragile and verbose. We’re fixing that by pushing the limits of model reasoning .
Each paper informs a product—whether it’s our InsurancGPT copilot, our custom GenAI solutions, or low-resource language models.
Exploring Next Token Prediction in Theory of Mind (ToM) Tasks: Comparative Experiments with GPT-2 and LLaMA-2 AI Models
Abstract
Language models have significantly advanced in their ability to generate coherent text and predict subsequent tokens based on given prompts. This study systematically compares the next-token prediction performance of two widely recognized models: OpenAI’s GPT-2 and Meta’s Llama-2-7b-chat-hf in Theory of Mind (ToM) tasks. To rigorously assess their capabilities, we constructed a diverse dataset using 10 short stories sourced from the Explore ToM Dataset. We enhanced these stories by programmatically inserting additional sentences (referred to as infills) using GPT-4, creating multiple variations to introduce varying levels of contextual complexity. This approach allows us to examine how increasing context influences model performance. We evaluate model behavior under different temperature settings (0.01, 0.5, 1.0, and 2.0) and test their ability to predict the next token across three distinct reasoning levels. Zero-order reasoning involves state tracking, which may probe either the current state (ground truth) or prior states (memory). Firstorder reasoning refers to understanding someone’s mental state (e.g., “Does Anne know the apple is salted?”). Second-order reasoning introduces an additional level of recursion in mental state tracking (e.g., “Does Anne think that Charles knows the apple is salted?”). Our findings reveal that increasing the number of infill sentences slightly reduces prediction accuracy, as added context introduces complexity and ambiguity. Llama-2 consistently outperforms GPT-2 in accuracy, particularly at lower temperatures, where it exhibits higher confidence in selecting the most probable next token. As question complexity increases, model responses diverge significantly. Notably, GPT-2 and Llama-2 exhibit greater response diversity in first and second-order reasoning tasks. These insights highlight how model architecture, temperature, and context affect next-token prediction, enhancing understanding of language model capabilities and limitations.
Other White Papers
Frequently Asked Questions
InsurancGPT is a private, agentic AI platform purpose-built for the insurance industry. Developed by Enkefalos and powered by the GenAI Foundry control plane, it delivers secure, explainable AI across the core workflows that drive insurance operations: underwriting, claims management, document processing, compliance, and analytics.
Unlike generic AI tools adapted for insurance, InsurancGPT is insurance-native. It understands the language, logic, and regulatory requirements of insurance workflows from the ground up. Every output is traceable to its source, every decision is auditable, and every deployment runs within the insurer's own infrastructure, ensuring full data sovereignty and compliance.
InsurancGPT is organized into six specialized products: InsureAssist, DocuSure, UnderwriteIQ, ClaimFlow, InsightEdge, and AutoLens. Each can be deployed as part of the full platform or independently.
AI solutions for insurance companies are purpose-built platforms that apply artificial intelligence to core operational workflows. Effective solutions are trained on insurance data and governed by insurance logic.
InsurancGPT delivers six AI solutions:
- InsureAssist: Context-aware AI assistant for employees and agents.
- DocuSure: Document intelligence with page-level source traceability.
- UnderwriteIQ: AI-driven underwriting workflows and risk assessment.
- ClaimFlow: Intelligent claims automation and fraud detection.
- InsightEdge: Role-based analytics from natural language prompts.
- AutoLens: Computer vision assessment of accident photos.
ClaimFlow is an AI-native claims management product that delivers intelligent, end-to-end claims automation. It covers every stage from first notice of loss (FNOL) through settlement.
ClaimFlow works through five core capabilities: configurable workflows, automated data validation, AI-powered fraud detection, embedded compliance checks, and full decision auditability. It delivers 45x faster claims triage and a 60% reduction in loss run processing time.
AI transforms claims management by automating intake, validation, triage, fraud detection, and compliance checking. It ensures faster decisions and reduced leakage while maintaining a fully auditable record.
Key stages include structured data capture at FNOL, automated data enrichment, AI-driven prioritization based on risk, and pattern recognition to identify high-risk anomalies before settlement.
FNOL (First Notice of Loss) is the initial report of a loss event. AI automates this by replacing manual processes with structured digital intake, real-time data validation, and automated exception handling.
With ClaimFlow, AI-powered FNOL automation reduces the time from loss event to active claims handling from days to minutes, scoring claims by complexity and routing them to the appropriate handler instantly.
AI automates data extraction and normalization, risk assessment, and compliance checks. This reduces submission-to-decision cycle time by 72%.
UnderwriteIQ capabilities include: automated risk assessment against guidelines, 90% improvement in SOV validation quality, rapid loss run processing, and continuous learning from underwriter decisions.
Yes. AI supports the binding process by automating pre-bind validation steps while human underwriters retain final authority. It ensures that by the time a quote reaches the binding stage, all compliance checks and risk validations are complete and documented.
AI moves beyond simple rules into intelligent systems that learn from human decisions. InsurancGPT delivers improvements across speed (72% faster cycle times), accuracy (90% SOV quality improvement), and compliance consistency, while keeping every decision explainable and reversible.
AI works by automating data-intensive workflows and augmenting human decision-making with evidence-backed recommendations. It applies to underwriting (risk assessment), claims (fraud detection), documents (traceability), analytics (real-time insights), and visual damage (computer vision).
The governing principle: every output is explainable, every decision is traceable, and human oversight is maintained throughout the full insurance value chain.
AI only matters when it creates measurable outcomes.
We align technology, governance, and economics to deliver value that holds up under scrutiny.