Mindcapita working paper · v1.0 · August 2026 · Not yet validated - see section 9
Organisations instrument their outcomes in fine detail while the psychological capacity that produces those outcomes goes largely unmeasured. This paper describes a platform that measures psychological capital - the four state-like resources we call Direction, Confidence, Recovery and Outlook - through structured AI-led voice conversations with every employee, scored against behaviourally anchored facets and cross-checked by an independently developed questionnaire. We review the construct and its evidence base, derive the measurement design from the validity literature on language-based and interview-based assessment, describe the conversational agent, the evidence rules and the scoring model in detail, set out the role and the strict limits of the optional camera channel under the EU AI Act, document the data architecture that keeps every result anonymous and aggregated, and present the main use cases. We close with what the method does not yet claim and the calibration study designed to earn those claims.
The psychological state of a workforce is a cost and performance factor of macroeconomic size. The World Health Organization attributes twelve billion lost working days per year to depression and anxiety, at an estimated cost of one trillion US dollars in lost productivity[35]. For the United States, Goh, Pfeffer and Zenios associate more than 120,000 deaths per year and roughly five to eight percent of annual healthcare costs with how companies manage their workforces[34], and a systematic review of cost-of-illness studies finds that 70 to 90 percent of the societal cost of work-related stress consists of productivity losses, not medical bills[33].
The standard instrument for looking at this factor is the annual engagement survey, and its limits are documented rather than anecdotal. Single-source self-report questionnaires carry systematic common method bias[26]; respondents underreport what is socially undesirable and overreport what is desirable, with the distortion depending on who asks and how[27]; and a meta-analysis of 2,037 organisational surveys with over 1.2 million respondents shows response rates around 60 percent for non-managerial employees, 37 percent for executives, and a decline over time that is only masked by ever heavier use of reminders and incentives[28].
The platform described here takes a different route: a structured, semi-standardised conversation with every employee, led by a voice-capable AI agent, evaluated against a published psychological construct under explicit evidence rules, and reported only as anonymous group values. The design goals are scientific traceability, radical anonymity, and results within a day rather than a quarter.
Psychological capital (PsyCap) originates in positive organizational behavior, defined by Luthans as the study and application of positively oriented psychological capacities that can be measured, developed and managed for performance improvement[1]. Four resources meet these criteria - hope, self-efficacy, resilience and optimism - and form a higher-order construct that predicts outcomes beyond each single component[2][3][4]. Unlike stable personality traits, PsyCap is state-like: in the founding study its four-week test-retest stability was 0.52, against 0.76 for conscientiousness and 0.87 for core self-evaluations[3]. It is malleable enough to develop and to erode - which is exactly what makes it worth measuring repeatedly.
The first meta-analysis across 51 samples with 12,567 employees found consistent positive relations of PsyCap with job satisfaction, commitment, well-being and performance, and negative relations with cynicism, turnover intentions, stress and deviance[5]. The most comprehensive synthesis to date - 244 studies, 254 samples, 96,416 employees - reports corrected correlations of .425 with supervisor-rated job performance, -.359 with turnover intentions and -.551 with burnout[6]. We quote the supervisor-rated figure deliberately: the far larger coefficients in that literature (work engagement .712, job satisfaction .683) are relations between self-reported attitudes and carry common method inflation - the same synthesis shows the engagement link dropping from .718 in cross-sectional to .610 in longitudinal designs[6].
Adjacent evidence at the business-unit level points the same way. The peer-reviewed meta-analysis by Harter, Schmidt and Hayes across 7,939 business units found generalisable relationships between unit-level engagement and customer satisfaction, productivity, profit and turnover[29]; Gallup's 2020 update - an industry meta-analysis of 112,312 units, to be read as such - reports a true-score correlation of .49 with composite performance and median top-versus-bottom-quartile differences of 23 percent in profitability and 81 percent in absenteeism[30]. These are cross-sectional differences, not causal effects, and we report them as such.
PsyCap responds to intervention. In a randomised pilot with 242 participants, a focused two-and-a-half hour intervention raised PsyCap with an effect size of d = 0.40 against d = 0.04 in the control group[9]; a randomised web-based training replicated the effect online[8]. A pre-registered meta-analysis of 41 controlled trials with 3,911 participants puts the honest average at d = 0.34 overall and d = 0.26 for the PsyCap composite - real, and modest[10]. Longer-term follow-ups support sustained gains for hope, resilience and optimism, but not for self-efficacy[11]. For an organisation this means: the measured quantity can be moved, and claims of transformation from single short trainings are not covered by the evidence.
The four resources are the ones psychological capital research knows as hope, self-efficacy, resilience and optimism - the HERO model[4]. We name them Direction, Confidence, Recovery and Outlook, because a manager reading a team report should not have to translate a psychology term first. The renaming is presentational only: each dimension keeps its research tradition, named beside it below, and carries the original construct as its identifier throughout the data model, so results stay comparable with the literature. Each dimension is operationalised through three facets - twelve in total. A facet is the unit that is actually rated; dimensions and the overall score are derived from facets, never rated directly.
Clear goals of one's own, several ways to reach them, and the energy currently going into them. Snyder's hope theory separates agency, the goal-directed energy, from pathways, the ability to find routes and alternatives. Both are needed; one without the other does not carry.[17]
The trust to take on demanding work, and to stand behind an own position when it is contested. Bandura showed that mastery experiences are the strongest source of self-efficacy; voicing a dissenting position is a distinct behaviour, studied by Van Dyne and LePine, with Edmondson describing the team conditions under which it appears.[18][61][19][20]
Getting back to work after a setback, staying in control under pressure, and drawing on support. Lazarus and Folkman modelled how load is appraised and coped with; Cohen and Wills showed that social support buffers stress rather than merely accompanying it, which is why using support is measured as its own facet. Masten described resilience as ordinary adaptation, not a rare gift.[21][22][25]
A realistic, positive view ahead, and a differentiated way of explaining setbacks - not positive thinking. Peterson, Seligman and colleagues distinguished explanatory styles: setbacks attributed to internal, stable and global causes versus situational, temporary and specific ones. Folkman and Moskowitz described positive reappraisal, finding genuine value in a hard experience afterwards.[23][24]
The standard PsyCap instrument, the PCQ-24, is psychometrically solid at the composite level (alpha .88 to .89 across the founding samples) but weaker underneath: the resilience and optimism subscales repeatedly fall below the conventional .70 threshold[3], a critical review catalogues further conceptual and measurement shortcomings[12], and in the largest cross-cultural test - 56,363 employees across twelve countries - only a reduced nine-item version without the optimism items survived metric invariance[13]. The openly published CPC-12 and its revision CPC-12R show sound psychometrics (composite alpha .87 to .91, measurement invariance across US and German samples)[14][15][16] and serve in our design as the external criterion for calibration - not as an item source. Our questionnaire items are written independently from the facet definitions; no PCQ or CPC item is used.
That psychological constructs can be measured from natural language is an established research programme, from dictionary-based text analysis[40] through open-vocabulary models that outperform fixed dictionaries[41] to full psychometric validation: language-based personality assessments converge with questionnaires at r = .38 on average, agree with self-reports as strongly as human informants do, and show six-month test-retest stability of .70[42]. Against external, non-self-report criteria the signal holds: everyday language predicted future depression diagnoses in medical records with an AUC of .72 in the half year before diagnosis[43].
Modern language models raise the ceiling of that tradition. Across 47,925 annotated texts in twelve languages, GPT-class models tracked human raters at r = .74 to .77 for sentiment against .30 for dictionary methods, with rerun reliability of kappa .93 to .99[48]; their annotation agreement can exceed that of trained human annotators[49]; zero-shot trait inference from text approaches supervised models[50]; and rubric-based automated scoring improves when the rating scale is explicit[51]. The field's own position papers insist on the caveat we adopt in section 9: capability is demonstrated, validity per instrument is not - it must be earned per construct, per language and per population[47].
Eighty-five years of selection research give structured interviews a validity of .51 against .38 for unstructured ones[46]. The closest structural precedent to our design - machine-scored automated interviews - shows convergence of .38 to .42 when models are trained against observer-style ratings, and near-zero validity when trained against self-reports[45]; automated personality perception before that operated at modest accuracy[44]. The design consequences run through this whole paper: a fixed conversation protocol, observable behaviour as the rating target, and every facet scored against explicit behavioural anchors rather than free impression.
The primary instrument is a semi-structured voice conversation of typically fifteen to twenty minutes, led by a realtime speech agent. The conversation follows a fixed agenda - arriving, a deliberate stock-take of the person's everyday work that later serves as anchoring material, opportunistic coverage of the twelve facets in free order, and a proper close - under explicit conversational craft rules: one question per turn, always bridging from what was just said, episodes over self-descriptions, no leading questions, no interrogation, no psychologising. The agent mirrors the participant's language and can conduct the conversation in several languages; voice, pace and turn-taking eagerness are configurable per campaign.
Two auxiliary components steer without being visible. A coverage monitor - a separate model call under a strict schema - tracks which facets are already evidenced, counting a facet as covered only on a concrete episode or clear specific statement, never on a self-label alone. A prompter channel feeds the agent the next open facet, a suggested segue and the remaining time; the agent must reformulate rather than recite. Interrupted conversations resume across sessions: the transcript is re-scored on re-entry, only genuinely open facets return to the agenda, and a structured resume note - hard-limited in content, with special categories of data under GDPR Article 9 explicitly excluded - restores conversational continuity without restoring risk. The agent is not covert: asked what is being evaluated, it explains honestly in plain language.
The agent does not judge its own completeness. That decision sits with the coverage monitor, which runs beside the conversation on its own model and returns, for each of the twelve facets, whether the transcript so far carries usable evidence. Its cadence adapts to how much is left: roughly every 25 seconds while eight or more facets are open, every 40 seconds below that, every 60 seconds in the final stretch. Statuses only ever move forward, so a later ambiguity cannot erase evidence already given.
Only when every facet is marked covered does the monitor raise its wrap-up flag. The agent then receives one more prompter message - close the conversation properly, thank the person, say goodbye, and only then call the closing tool. The closing is therefore always spoken before it is executed; the session stays open while the farewell is still playing. Two tools exist for this, one that ends the conversation and one that pauses it, and the agent is instructed to speak first in both cases.
Time is a guard rail, not a metronome. After thirty minutes in one sitting the agent offers to continue another day and pauses; after sixty minutes of cumulative talk time across all sittings it closes with whatever evidence exists. Participants can pause or end at any point themselves, and a paused conversation resumes through the same personal link.
A last gate sits on the participant's own ending. If someone finishes early, the client re-checks coverage - counting a partially evidenced facet as half - and below 90 percent the conversation is not completed at all: the progress is stored and the state returns to paused. A half-measured conversation never reaches the aggregate.
Scoring is strictly bottom-up and rule-bound. Each of the twelve facets carries six behaviourally anchored scale steps, and the scoring model is instructed to choose the step whose description fits the evidence, not the impression in between. The evidence rules are absolute: behaviour beats labels - a concrete episode is strong evidence, a self-description weak; no evidence, no score - a facet without solid material stays unrated and is never filled with an average; contradictions favour the episode; and every rating must cite short verbatim quotes from the transcript as its evidence. Each rating carries a confidence derived from evidence density. A second, independent consistency pass re-checks every rating against its quoted evidence and may correct it by at most one step, with the correction and its reason persisted.
A dimension is the confidence-weighted mean of its rated facets; a dimension without evidence is omitted, not zeroed. The four dimensions form the overall score on the six-point scale. Where the live monitor and the scorer disagree about a facet twice, the facet is dropped from the agenda and stays unrated - not measurable beats a guess. When and how a conversation ends is described in section 5.2.
Alongside the rating, the conversation collects something the rating cannot supply: what the person says would make their work easier. The agent asks for this two or three times, in passing and always attached to a topic just discussed, never as a list and never with a promise that anything will follow. A separate extraction pass - deliberately separate, so it can neither see nor influence the ratings - turns the answers into short phrases without names, teams, projects or quoted wording, each assigned to one dimension.
These levers are never part of a score. A wish is not evidence of a state, and mixing the two would corrupt the measurement in the direction of what people want rather than what is. They are aggregated per dimension across a campaign, counted per person rather than per mention so one talkative participant cannot make a theme look widespread, and reported only from five people upwards - the same floor that governs every other figure. Below it, a lever is one person's sentence and stays invisible. Where a theme clears the floor it is shown ahead of the generic recommendation, because it is the only part of the advice that knows this particular organisation.
A twelve-item questionnaire - one independently written item per facet, three reverse-coded against acquiescence, on the same six-point scale - can follow the conversation after a configurable delay. It exists as a cross-check, not a second truth: the current person-level score weights the latest conversation at 0.65 and the latest questionnaire at 0.35, and where group means from the two instruments diverge by more than one scale point, the divergence itself is surfaced as a finding. These weights are starting values, disclosed as such, and belong to the calibration described in section 9.
A language model is an instrument, and instruments drift. Semantically irrelevant prompt-format changes alone can swing model accuracy by tens of points[52], so prompts are fixed and versioned, scoring runs at temperature zero under strict output schemas, and every conversation records the exact model set that produced it, keeping scores comparable and re-scorable across model generations. Audits of language models rating people have found measurable demographic disparities[53] and Western-centred emotion norms even in multilingual models[54]; fairness and per-language audits are therefore part of the validation programme, not an afterthought.
The conversation can be accompanied by the participant's camera. This channel is optional, off by default, governed by its own separate consent, and processed entirely on the participant's device: a local face-landmark model computes a single momentary valence tendency from a handful of facial movement signals. No video and no frame-level data ever leave the browser - the connection to the speech model carries audio only - and the only artefact that is stored is one aggregate summary per conversation, reported solely at group level under the same five-person minimum as every other figure. The camera signal is excluded from measurement by architecture: no scoring path reads it.
The scientific case against inferring emotion from faces is strong and we adopt it. The most thorough review of the evidence concludes that facial configurations map to emotion categories with limited reliability, no specificity and limited generalisability - the meta-analytic association between proposed configurations and their emotions averages r = .31, and the proposed configuration appears in only about 22 percent of matching emotional events[56]. The engineering numbers point the same way: human annotators agree on in-the-wild emotion labels only 60.7 percent of the time[58], recognition accuracy drops from near-perfect on posed laboratory data to 48 to 75 percent in the wild[57], and multimodal fusion helps far less on natural data than on acted data[59]. Signals of this reliability class have no place in a psychometric score. Facial movement description as such[55] remains a legitimate research field; treating it as an emotion fingerprint does not.
Regulation (EU) 2024/1689 prohibits, since February 2025, the use of AI systems to infer emotions of natural persons in the workplace, with exceptions only for medical or safety purposes; its recital 44 cites the same three shortcomings - limited reliability, lack of specificity, limited generalisability - as the scientific review above[60]. We treat this prohibition as binding design guidance. The platform does not infer emotion categories, does not build affect profiles, does not use the camera channel in any measurement, and operates it only on-device, consent-gated and fully optional; organisations can run the platform with the camera channel disabled entirely. Whether even an aggregate valence summary falls within the scope of Article 5(1)(f) is a question we treat conservatively and under continuing legal review - the measurement described in this paper never depends on it.
Identity and content are separated by design: the invitation system knows who was invited, the evaluation pipeline works on conversation content, and the two meet only in aggregate. Every reported figure - campaign results, organisation-wide analysis, every attribute-filtered slice, the questionnaire cross-check and the camera summary - is suppressed below a minimum group of five completed participations, re-applied after every filter so that slicing can never isolate individuals. No transcript and no individual score is visible to the employer in any interface or API, at any role level. Participation begins with explicit consent that states exactly this; everyday language is sensitive data - it can predict clinical conditions[43] - and the architecture treats that as a constraint, not a footnote. Auxiliary artefacts that carry content between sessions are hard-limited in scope and explicitly exclude special categories of data under GDPR Article 9.
The same architecture is built to measure whole organisations quickly. Admission control caps concurrent live conversations against the capacity of a pool of realtime-model credentials; invitation pacing spreads outreach at a configurable rate; interrupted sessions survive network loss with continuously synchronised transcripts and can always resume; and campaign configuration freezes at go-live so that a running measurement cannot be silently altered. A measurement across thousands of employees returns its first aggregate results the same day.
When market share and thought leadership slip while product and pricing reviews find no cause, the psychological capacity of the organisation is a candidate explanation that standard controlling cannot see: PsyCap correlates negatively with burnout (-.551) and positively with rated performance (.425)[6], and unit-level psychological state tracks hard business outcomes[29][30]. A cross-organisational measurement locates the units where the drift begins - before it reaches the numbers.
Financial, technical and legal audits leave the acquired workforce unexamined, while the academic record shows persistently high M&A failure rates with human and cultural factors as the under-examined explanation[36] - Christensen and colleagues put the failure rate at 70 to 90 percent, a practitioner claim we cite as such[39]. Experimentally, merging two groups with incompatible conventions degrades performance, and members misattribute the loss to the other side's people rather than the clash itself[37]; in field data, measured cultural fit predicts deal success[38]. A PsyCap screening of the target makes its psychological capacity part of the due diligence record before signing rather than a post-closing discovery.
The best predictors of leaving are psychological states, not file data: quit intention correlates with actual turnover at .38, organisational commitment at -.23 and job satisfaction at -.19, against -.09 for age and .05 for education[31]; PsyCap itself relates to turnover intentions at -.359[6]. Quiet disengagement precedes the visible exit and has entered the scholarly literature as a construct of its own[32]. Repeated group-level measurement shows which teams are quietly disengaging while intervention is still possible - exit interviews only explain it afterwards.
The measure has not yet been validated, and we treat that as the central open item of this paper. Capability precedents from the literature do not transfer validity to an individual instrument[47]; it must be demonstrated. The calibration study in progress follows the validation template of the language-based assessment literature[42][45]: convergent validity against CPC-12R[15][16] as a free, published external criterion; agreement with trained human raters scoring the same transcripts; retest stability; discriminant validity between facets; and criterion relations over time. The benchmarks it has to meet are public: convergent correlations around .38 to .42 and rerun reliability above .9 are the standards set by the closest prior work[42][45][48].
Until then we make no claim about accuracy, reliability or predictive power, and we publish no figures on them. Known limits beyond validation: the PsyCap outcome literature is largely cross-sectional and its largest coefficients carry common method inflation[6]; PsyCap relations vary by country and culture, so any benchmark must be population-sensitive[6][13]; language models can encode demographic and cultural bias that must be audited per language, not assumed away[53][54]; and the internal weighting constants of the scoring model - the source weights, the risk threshold, the confidence cut-off - are disclosed starting values, not calibrated parameters. Intervention effects on PsyCap, where an organisation acts on results, should be expected in the d = 0.3 range shown by controlled trials[10], not in the range of promotional claims.