Zurich has the Leadership Expectations. This blueprint makes them breathe: one engine that drives success profiles, assessment, triangulated reporting, development and twelve months of measurable movement, for every development centre going forward. It is the architecture now in delivery for the CIO Development Centre in September 2026, generalised for enterprise use, with nothing withheld.
Every element is clickable. The cycle is the point: step six feeds step one, which is what makes the framework a living system rather than a document.
Most competency frameworks describe leadership; few get to run anything. This design puts the Leadership Expectations in charge of every stage: the profile is written in its language, evidence is mapped to its themes, raters answer against its indicators, reports are organised by its dimensions, and measurement tracks movement on its terms. One language from calibration to reassessment.
Nothing in this blueprint is proprietary mystery. Every design decision, every declared limit and every judgement rule is stated so that Group Talent can govern the model with full sight of it. And the cycle is deliberately built so that Zurich carries it: managers, sponsors and HR partners run the development year and the reassessment, and Zurich's own accredited practitioners run the instruments and 360s in house from day one. External involvement concentrates only where independence genuinely adds value: the design and build, the integration judgement at the centre, and light touch QA.
Everything here is drawn from a build that exists: the CIO Development Centre architecture, its five exercises, its instrument battery, its 360 modelled on the Leadership Expectations, its triangulated report and its manager and HR outputs. The blueprint generalises that build so any unit can run the same cycle against its own success profile.
The Leadership Expectations organise what Zurich means by leadership. A development centre needs one thing more: a decision about which of those expectations matter most for a specific population, at a specific altitude, right now. That decision is the success profile, and writing it is the first act of every centre.
Dimension definitions and theme names are Zurich's, verbatim. The domain tag on each theme (Thinking, Social, Self, Drive) is the assessment layer added beneath the framework: it groups themes by the capability they draw on, which is what makes evidence mapping and triangulation systematic in steps two and three. Learning Agility carries an integrating tag: its evidence feeds the transition readiness layer described in step two.
Every unit that commissions a centre completes one structured calibration session, roughly ninety minutes with the sponsor and Group Talent. It converts the enterprise framework into that unit's standard, in a format that is identical across the organisation. The format matters as much as the content: because every profile is captured the same way, Group Talent can lay profiles side by side and read the enterprise's competency demand as one map.
The unit states its three business priorities for the profile's horizon. Everything selected afterwards must serve one of them.
Up to eight reflecting themes are selected as critical at the target altitude, each tagged to a priority. The cap is deliberate: a profile that emphasises everything emphasises nothing.
Weighting is distributed across the selected themes, and any dimension scoped out of exercise assessment is declared now, not discovered at write up.
A fixed global core set protects the common leadership culture, and Group Talent may add up to two stretch expectations per profile.
Worked instance: the CIO Leader Success Profile selected twenty four themes across six dimensions, scoped Bold Advocacy out of exercise assessment while keeping it visible through the 360, and set the benchmark at business leader altitude.
Not every selected theme can be assessed to the same depth in a one or two day event, and pretending otherwise is where centre credibility usually dies. The profile therefore places every theme in one of three tiers, published before the event. Declared limits are a trust asset: participants, raters and sponsors all see in advance exactly what the centre claims to measure and what it does not.
Themes with at least three independent evidence sources, including direct behavioural observation. These carry ratings against the profile standard.
Themes rated with a stated constraint, for example evidence led by the 360 and instruments rather than by exercises. The limit is printed wherever the theme is reported, on every page, in every document.
Themes a centre cannot rate honestly, reported in words through the 360 alone, and never carrying a rating marker. Saying so plainly is the design, not a shortfall.
Every rating in the entire cycle is made against one fixed benchmark: the profile's declared altitude within the Leadership Expectations, business leader or people leader. Never against the other participants. Cohort ranges appear in reporting for context, but no one's result moves because of who else was in the room.
No single source can measure a leader. Observation shows what someone does; instruments show what they have and what they want; raters show what it is like to be on the receiving end. The design starts from what each source can honestly reach, then combines them so that every rated theme rests on at least three independent lenses.
A bespoke fictional company gives every exercise one coherent world: group strategy work, difficult one to ones, stakeholder negotiation, committee decision making and a timed case study, selected to cover the profile's critical themes. A two day format is a repeated measures design: revisiting a decision on day two turns adaptation itself into observable evidence. Observers follow a stream then code discipline, capturing behaviour live and coding it afterwards, so no rating rests on memory or on first impressions.
Logiks Advanced, normed against a senior manager group. Reasoning capacity anchors the Thinking domain and feeds the ability strand of transition readiness: how quickly a leader can take on genuinely new material.
Hogan HPI for day to day style, HDS for what emerges under pressure, MVPI for what fuels the leader. Together they explain the behaviour the exercises surface, and the HDS in particular powers the let go conversation in step five: derailers are strengths overused, and naming them that way is what makes them workable.
Raters answer the Leadership Expectations' own reduced indicators for the target altitude, by rater group: manager, peers, direct reports, stakeholders. Because the 360 speaks the same language as the exercises and the profile, its results triangulate directly instead of needing translation. It is also the only honest window into themes that two days cannot stage, which is what makes the commentary tier possible.
A competency based interview against the profile, useful where career evidence matters, where a population is small, or where a phase one and phase two design spreads assessment across a year.
Functional expertise judged by senior business observers in a separate lane, reported as structured commentary and never blended into behavioural ratings. Business judgement stays in the room; assessment judgement stays with the assessor team.
Calibration ends with a published matrix like this one, drawn here from the CIO build: which sources carry each dimension, where evidence is primary and where secondary, and the rule beneath it all, a minimum of three independent sources behind every rated dimension. Where structure limits coverage, the limit is declared in the profile, which is exactly where the tier system in step one comes from.
| Dimension | Group exercise | One to one | Stakeholder | Committee + follow up | Case study | Instruments | 360 |
|---|---|---|---|---|---|---|---|
| Winning Mindset | |||||||
| Enterprise Perspective | |||||||
| Strategic Execution | |||||||
| Relentless Drive | |||||||
| Interpersonal Effectiveness | |||||||
| Development Ownership |
Transition readiness is not scored as another competency. It is read as an integrating layer across all sources: ability (cognitive capacity), motivation (curiosity and learning values in the instruments) and application (adaptation actually observed between exercises and between day one and day two). It answers the one question the ratings cannot: how ready is this leader for the next transition, and how fast can this picture change. Step five shows how the lean into, let go pair is read through this lens, and why keeping it descriptive, separate from the competency ratings, is a deliberate design decision.
Reporting is organised competency by competency, never instrument by instrument. For each theme, three voices are laid side by side. Where they agree, the picture can be trusted. Where they diverge, the divergence is not averaged away: it is usually the most valuable finding on the page, and it is named.
What trained observers watched the leader actually do, captured live in the exercises, coded after the fact.
What the instruments say about wiring and capacity: style, pressure patterns, values, reasoning.
How nominated raters experience the leader over months and years: manager, peers, direct reports, stakeholders.
Three independent lenses converge on every theme. The toggle below drives this picture: watch the reputation line when the voices stop agreeing.
In the committee exercise, cut the option list from five to two within ten minutes and held the group to an explicit decision rule when new information arrived on day two.
Profile shows high standards of order and goal focus with strong commercial values: a leader wired to rank, sequence and finish.
All four rater groups place prioritisation among the highest rated indicators; direct reports describe always knowing what matters this quarter.
Ranked options decisively in every exercise; observers rated the behaviour clearly at standard.
Instruments add a caution: under pressure the profile shifts towards restless course changing, with excitement about the new crowding out the committed.
Direct reports rate the same indicators noticeably lower than peers do, describing priorities that change faster than the team can absorb.
Observers capture behaviour verbatim as it happens and code it against themes only afterwards. Separating capture from classification prevents the oldest assessment centre failure: seeing what you expected to see.
Different assessors observe each participant across exercises by design, and every rating is challenged and confirmed by the full team in a structured integration session before it stands.
Every evaluative passage in reporting is attributed: the observing assessor, the lead psychologist's integrated judgement, or pre authored wording selected by confirmed results. No evaluative sentence is machine written, and the reader can always tell.
A theme that entered the event with a declared limit keeps that limit through integration and into every document that reports it. Integration is where evidence is weighed, never where caveats quietly disappear.
Everything up to here is excellent process. This stage is professional judgement: weighing an observed behaviour against a pressure pattern against a rater split, deciding what the evidence can honestly carry, and writing the sentence a leader will act on. It is the stage that requires registered assessors, calibration hours together, and the independence to say that the evidence does not support a rating, even when a neater answer would be more convenient.
The same triangulated evidence is written three times, for three readers, and the discipline is what does not travel between them. The participant owns the fullest picture. The manager gets what sponsorship needs. HR gets the bench, never the verbatims.
Twelve pages, written in the second person, organised by the profile's dimensions rather than by instrument. Each dimension page triangulates the three voices, shows the rater group breakdown with anonymised verbatim comments, states any declared limit, and points forward: where this goes next. It is written as a working document, a mirror with a direction, and the margins belong to the reader.
A single landscape page for the sponsorship conversation: the four domain read, band bars triangulated across sources, the one gating gap if the profile has one, a horizon sentence, and the three sponsorship moves that would make the biggest difference this year.
Readiness and trajectory content lives here only. It never appears in the participant document, and the participant is told the summary exists and may see its substance on request.
The bench in one view: a readiness map (ready now, one to two years, develop longer term), a cohort theme heatmap against the profile, and the development themes the cohort shares, which is where group interventions earn their budget.
Aggregated and role focused: no verbatim rater comments, no personality subscale data. What HR needs to plan with, nothing it does not.
A report that ends at page twelve is an event. Each output is designed as the opening move of the twelve month cycle in step six, and each spin off below exists because a specific reader needs a specific next step.
This is the blueprint's development philosophy, and it is deliberately not the classic strengths and development areas read. The two approaches use the same evidence and produce very different conversations. The difference is worth being precise about, because it changes what leaders actually do on Monday.
The traditional read treats the profile as a checklist: the highest bars are labelled strengths, the lowest become the development plan, and growth is assumed to mean closing gaps towards a uniform template. At senior level this quietly fails three times over.
It buys improvement where it is most expensive. A seasoned leader's lowest scores usually reflect stable disposition. Moving them from below standard to nearly standard consumes a year of effort for marginal, often invisible, return.
It ignores where impact actually comes from. Senior contribution is carried by two or three distinctive capabilities, not by an absence of weaknesses. A plan that spends nothing on the strengths leaves the leader's real leverage undeveloped.
It cannot see overused strengths. The most common senior derailment is not a missing capability but a career making capability applied too hard, too often, at the wrong altitude. On a gap read, that pattern scores well and goes unexamined.
Lean into, let go reads the same profile as a question about movement between altitudes: what got this leader here, and what will get them there.
Lean into names a confirmed strength with headroom: a capability the evidence triangulates, that the next altitude will reward disproportionately, and that the leader is currently using at a fraction of its reach. The development act is deliberate, expanded, more visible use.
Let go names a success habit to retire: a behaviour that built the career at a previous altitude and now caps it. Very often it is the shadow of a lean into strength, which is why the pressure profile in the instruments matters: derailers are strengths overused, and naming them that way makes them workable rather than shameful.
The pair is the point. One thing to expand, one thing to retire. Specific enough to be observable, which is what makes it rateable by others, which is what makes step six's measurement honest.
A worked illustration on a sample six dimension profile. Toggle between the two readings: the bars never change, the development conversation does.
Movement does not abolish standards. If a theme sits below the line the role requires, it is named as a gating gap, it goes to the manager conversation, and closing it joins the agenda. The discipline is that gating gaps must be earned by the evidence against the profile standard: a bar being lowest is not by itself a development need.
Competency ratings answer where the leader is now. Transition readiness answers how fast that picture can change, and it is the lens this whole page is read through: readiness signals tell you how much headroom a lean into really has, and how much a let go will cost the leader to retire. Keeping readiness descriptive, and separate from the competency scores, is what makes that read possible. Deriving one from the other would make both redundant, and mirrors the same logic as Zurich's own performance versus potential distinction.
Start from the strength and mean it: here is what the evidence says you are unusually good at, and here is where using more of it would matter most next year. Then the harder gift: here is the habit that built your career and is now taxing your team, and here is the pressure pattern that explains it. Agree one experiment for each, observable within ninety days. Close by checking the standard: nothing gating, or one thing gating, named plainly. That is the whole conversation, and leaders act on it because it takes their success seriously instead of treating them as a set of deficits.
A centre that cannot show what changed was an event, not an investment. The measurement design makes a deliberately narrow claim: visible movement on each leader's named lean into and let go behaviours, plus a bench that Group Talent can plan against. Narrow is what makes it true. Click any milestone.
Assessment against the profile, integration, and the three outputs from step four. Participant 360 detail is held back until the individual feedback session, so the first encounter with reputation data happens with an assessor in the room, not alone at a desk.
Each leader has their individual session and leaves with an agreed agenda: one lean into experiment, one let go experiment, any gating item, all observable within ninety days. Sponsors are briefed on the one pager and commit to their sponsorship moves. This month decides whether the year is real; it is resourced accordingly.
A short structured conversation per leader: which experiments ran, what they surfaced, what the agenda now says. Experiments that did not run are treated as data, not as failure: the blocker is usually the finding.
The original raters are re asked only the indicators behind each leader's named lean into and let go themes: eight to ten items, minutes to complete, high response rates because it is short and personal. Because let go targets are specific behaviours, raters can genuinely see whether they stopped, which is the measurement advantage of the whole approach. A sponsor pulse runs alongside: are the sponsorship moves happening?
The development agenda enters the year end conversation through the leader's own manager, in the framework's language. The ecosystem's quiet win: managers and the centre describing leadership with the same words.
Group Talent receives the year's evidence: pulse deltas, agenda completion, bench movement. Reassessment is then designed to be run in house. Each manager receives a leader specific reassessment kit, built externally once: a structured review conversation generated from that leader's agenda, observation prompts, a rating rubric against the profile, and a refreshed pulse 360. The manager runs it, the HR partner facilitates, and the external practice's role narrows to building the kits and quality assuring a sample of the results, with Zurich's own accredited practitioners running any instrument re runs. A targeted external mini centre is reserved for the few cases where a live promotion decision needs fully independent evidence. The findings feed the next calibration: profiles are refreshed with what the organisation now knows, which is the arrow that makes the map on page one a cycle.
Because every profile is calibrated in the same format and every centre reports in the same language, cohort data aggregates. Two or three centres in, Group Talent holds something no framework document can produce: a live heatmap of which Leadership Expectations the enterprise demands most, where the bench is deep, and where it is thin. The framework stops being a poster and becomes the organisation's talent operating data.
Illustrative renders with sample data, participants anonymised P1 to P10. Every deliverable speaks the framework's language, which is what lets them stack: individual pages roll into the manager view, manager views roll into the cohort dashboards, and the year's data rolls back into the next profile.
Everything above was built for the CIO Development Centre and runs in September. Because it is all written in the framework's own language, none of it is single use: the same report architecture, kits and dashboards can carry any unit's profile with a calibration session and a change of cover. Zurich's own accredited practitioners run the instruments and 360s, managers, sponsors, HR partners and Group Talent carry the year, and external involvement concentrates where independence adds most, in the design, the integration judgement at the centre itself, and light touch QA. One build for the CIO population, a blueprint for every population after it.