Skip to content

The 9-Box Grid for Succession Planning: GCC HR Guide

A 9-box grid maps performance against potential so you can spot future leaders before a resignation forces the question. How GCC HR teams use it.

Sayed Hussain AlmukhtarContent Writer, Lumofy
14 min read

A key employee resigns without warning, and there's no one ready to step into the role. A strong performer gets promoted into management, and six months later, it's clear they weren't ready. A 9-box grid is the tool built to prevent both: it plots every employee on performance and potential, sorting people into nine categories that show who's ready now, who needs development, and who's exactly where they should stay.

Used properly, it turns succession from a guess into a plan, which matters more than it sounds. SHRM's 2021 survey of 580 HR professionals found that 56% of organizations have no succession plan at all, most often because of limited time and resources, or because the business felt too small to need one (SHRM, 2021).

Gartner's numbers are just as stark: 72% of HR leaders say they struggle to close gaps in successor readiness, and only 38% of CHROs are confident they can actually deliver on their succession goals this year (Gartner, Abhinandan Sood, September 2025).

of organizations have no succession plan at all
56%
SHRM, 2021 survey of 580 HR professionals
of HR leaders struggle to close gaps in successor readiness
72%
Gartner, September 2025
of CHROs are confident they can deliver on their succession goals this year
38%
Gartner, September 2025

Most of that gap traces back to a missing tool, not missing talent. Nobody had a way to see who was actually ready before the moment forced the question.

Where did the 9-box grid come from?

The 9-box grid wasn't built to evaluate people. McKinsey developed it with General Electric in the early 1970s to help GE decide which business units deserved more investment and which didn't, plotting them on industry attractiveness against competitive strength (McKinsey & Company, "Enduring Ideas: The GE and McKinsey Nine-Box Matrix").

HR departments noticed the same logic worked for people and swapped the axes: competitive strength became individual performance, and market attractiveness became future potential. GE turned it into an annual ritual called Session C, a top-to-bottom leadership review that started under CEO Reg Jones and was sharpened further under his successor, Jack Welch. By the 2000s, some version of it was standard practice across most large HR functions.

What do performance and potential actually measure?

Performance is the easier axis. It's what someone actually delivered against clear goals over a defined period, backed by data most managers can already agree on.

Potential is harder, because it's a bet on the future rather than a record of the past. It usually shows up as ambition, how fast someone learns, and how well they adapt when the ground shifts, and none of those are easy to measure with precision. That's exactly why potential ratings drift toward bias faster than performance ratings do.

The Corporate Leadership Council's four-part test still holds up as a practical check:

  • Aspiration: genuine desire for more responsibility, not just visibility.
  • Ability: how someone handles new complexity, not just familiar problems.
  • Engagement: real investment in the business's actual goals.
  • Agility: how quickly a setback becomes a lesson instead of a grudge.

Why do ratings need calibration?

Picture two managers rating employees who do the same job. One gives a 4 out of 5. The other gives the same performance a 2.

The difference usually comes down to which manager's personal bar the rating happened to pass through, not which employee actually performed better.

That's what calibration sessions exist to fix: managers review ratings together, backed by real examples rather than gut feel, until similar performance gets a similar score regardless of who's grading it. Skip this step, and a 9-box grid measures management style, not talent.

Calibration sessions typically surface four recurring biases:

Get this right, and the grid stops reflecting whoever happened to write the review, and starts reflecting the people it's supposed to describe.

What are the nine boxes, and what does each one mean?

Once performance and potential are both calibrated, the grid itself is the easy part.

The 9-box grid in Lumofy's Potential Insights, showing all nine boxes with talent counts and recommended actions

High potential

  • High performance (Future Leader): the person whose departure would actually hurt. Real investment and retention effort belongs here.
  • Medium performance (Emerging Talent): needs harder assignments and focused coaching to close the gap between promise and output.
  • Low performance (Enigma): the one to slow down on. Low performance here is often a bad role fit or missing guidance, not missing ability.

Medium potential covers most of the org chart, and that's fine.

Medium potential

  • High performance (High Impact Performer): excelling right where they are, with no clear signal they're ready for more yet. The right move is recognition and a wider role, not a promotion.
  • Medium performance (Core Player): the backbone of the team. Protecting their stability matters more than rushing them anywhere.
  • Low performance (Dilemma): needs a direct conversation to find out whether the fix is training or a different role entirely.

Low potential doesn't mean low value.

Low potential

  • High performance (Trusted Expert): worth more staying deep in their specialty than getting pushed into management they don't want.
  • Medium performance (Solid Performer): doing fine, and doesn't need urgent attention.
  • Low performance (Underperformer): the one box that calls for a real improvement plan, or a hard conversation about whether the role still fits.

It's normal, and healthy, for Core Player to be the biggest box on the grid. An organization's value doesn't come from its stars alone; it comes just as much from the people who show up and hold steady.

Replacement planning vs. succession planning: what's the difference?

The two get used interchangeably, and that's a mistake, because they solve different problems.

Replacement planning kicks in after a role is already empty: a resignation, a termination, an emergency. The goal is to fill the gap fast, usually through external hiring or promoting whoever's closest and available.

Succession planning starts long before anyone leaves. It's deliberate preparation for roles where a sudden vacancy would actually hurt, built through development plans that get candidates ready ahead of need, not after it.

AspectReplacementSuccession
TimingAfter the role is already vacantLong before any vacancy
TriggerResignation, termination, emergencyDeliberate planning for critical roles
GoalFill the gap fastHave a ready candidate before the need arises
Typical methodExternal hire, or promote whoever's closestDevelop candidates in advance, on a real plan

The cost argument backs this up. A Wharton study of personnel data from a US investment banking division found that external hires were paid roughly 18 to 20% more than internal promotes doing the same job, scored lower on performance for their first two years, and left or were let go at meaningfully higher rates than people promoted from within (Wharton, Matthew Bidwell). Replacement is expensive even when it works.

Can a one-person HR team actually run this?

All of this can sound like it needs a talent management department to run. It doesn't. A single HR generalist can do it well, as long as they simplify the process and focus on what matters most.

What makes that possible comes down to four practices:

  • Mini calibration sessions. Pull the direct managers of critical roles into one focused session instead of reviewing every employee solo.
  • Standard questions, asked the same way of every candidate, so ratings don't depend on who happened to be asking.
  • A narrow starting scope: only the roles whose sudden loss would actually disrupt the business, not the whole company at once.
  • Everything written down, so decisions don't rely on anyone's memory six months later.

From there, the practical sequence runs in six steps:

  1. Identify the critical roles: the ones whose sudden vacancy would genuinely disrupt the business. Gartner's guidance is to flag roughly 10 to 15% of all roles this way, not to try to plan for everyone.
  2. Assess where current role-holders stand today, using whatever performance data already exists.
  3. Define what the role will actually require in a few years, not just today.
  4. Give candidates real stretch assignments and real coaching, not just a spot on a list.
  5. Build more than one credible successor per critical role, so losing one person doesn't send you back to zero.
  6. Review it on a rhythm: one full pass a year, with a lighter check-in every quarter.

None of this needs headcount. It needs a plan.

How does this connect to Saudization, Emiratisation and the rest of GCC localization?

Every Gulf state runs some version of a nationalization drive: Saudi Arabia's Nitaqat, the UAE's Nafis, and similar pushes in Qatar, Oman, Bahrain and Kuwait. All of them share one goal: more citizens working in the private sector, not just the public one.

Saudi Arabia's numbers show what's actually at stake. Saudi nationals working in the private sector passed 2.27 million by the end of 2024, up from around 1.8 million in 2021, according to the Ministry of Human Resources and Social Development and the General Authority for Statistics (MHRSD, GASTAT). The harder question sits underneath that number: how many of those roles are genuinely directive or leadership positions, rather than execution-only ones?

That's the real fork between two ways of running a nationalization program. One treats it as a compliance exercise: hit the Nitaqat color band, avoid the penalties, move on. The other treats it as a talent strategy.

Nitaqat's own mechanics already lean toward the second version more than most people realize. Its wage-weighted calculation counts a Saudi employee earning above SAR 4,000 a month as a full point toward the ratio, one earning between SAR 3,000 and SAR 4,000 as half a point, and one below SAR 3,000 as effectively zero, with extra weight given to engineering, medical and leadership roles (Motaded, 2026; confirmed independently by Vialto Partners). The mechanism already rewards better roles, not just more headcount.

MHRSD approved a new phase of Nitaqat Al-Mutawwir, first launched in 2021, aiming to localize more than 340,000 additional private-sector jobs over three years (Argaam, January 2026). That phase went live on April 26, 2026: fixed workforce bands gave way to a logarithmic scaling formula, and compliance moved from periodic review to real-time monitoring through the Qiwa platform (EY; Vialto Partners). Saudi commentary on the rollout described it in blunt terms: a shift "from quantitative compliance to qualitative impact" (Al Weeam, April 29, 2026).

A 9-box grid is what makes "better roles, not just more headcount" operational instead of aspirational. Used on national talent specifically, it identifies who already combines strong performance with real potential, then builds a development path toward running whole functions, not just staffing them. That's also the direct ask behind Saudi Vision 2030's Human Capability Development Program: national talent ready for leadership and specialist roles, not headcount on a roster (Vision 2030, Human Capability Development Program).

Why does the grid need to live in one system, not five?

Here's where most 9-box efforts quietly fail: not at the grid itself, but at what happens around it.

The pattern is familiar. Goals live in one spreadsheet, ratings get discussed in a meeting nobody wrote up, and the development plan ends up in a document that's forgotten within weeks.

Each piece might be done well. Scattered across five places, none of it adds up to a system anyone can actually run succession off of.

The fix isn't a better spreadsheet, and it isn't treating the 9-box exercise as a side project layered on top of the performance cycle everyone's already running. The Potential Insights tab in Lumofy's Performance Insights dashboard, for one, generates the grid straight from a completed review cycle: it places each person automatically using their performance score and assessed potential, sorting them into nine segments (Future Leader, Emerging Talent, Enigma, High Impact Performer, Core Player, Dilemma, Trusted Expert, Solid Performer and Underperformer), then surfaces AI-suggested next steps and candidate future roles for each person, viewable by admins and super admins inside any active or completed cycle. Nobody has to reassemble the picture from three different files, because the review cycle already had the data.

Lumofy's AI Justification panel, showing an employee's 9-box position, performance and potential ratings, and the AI-generated reasoning behind them

That's the actual case for running this inside one system rather than across several: not that the spreadsheet version is impossible, but that it depends entirely on someone maintaining it by hand, and that's exactly the kind of thing that quietly stops happening by month four.

When is a 9-box grid the wrong tool?

None of this is worth doing on a team of five.

Below a certain size, the grid stops producing signal and starts producing noise: every rating gets swayed by whatever happened last week, and a manager who already knows a five-person team well doesn't learn anything new from formalizing what they already know. Guidance synthesized from Gartner's and SHRM's succession-planning frameworks puts the useful floor at roughly 20 or more comparable people spread across more than one manager, since that's roughly where real patterns start to separate from individual impressions. A few sources put the floor as low as fifteen, but the signal only really firms up past twenty.

FAQ

Most practitioners put the useful floor at roughly 20 or more people, and accuracy improves as the group grows and spans more than one manager. Below that, a simpler three-category split (invest, stable, needs intervention) does the job without forcing categories that don't have enough people to fill them meaningfully.

A calibration session is a meeting where managers review their ratings together instead of scoring in isolation, comparing similar cases side by side against real evidence rather than gut feel. The goal is a shared standard across the team, not a shared average.

Forced ranking sets the outcome first: it requires a fixed share of top, middle and bottom ratings regardless of actual performance. Calibration sets the process instead, aiming for a fair result through shared evidence, whatever shape that result turns out to take.

Run a full calibration session once a year, covering every critical role across the business. Add lighter, faster check-ins every quarter to track how ratings are moving without repeating the entire session each time. Skipping the annual pass and relying only on quarterly check-ins is usually how ratings drift back to being manager-dependent.

Usually not by name. Telling someone they're a "Dilemma" or a "Core Player" turns an internal planning label into a verdict they didn't ask for. What they need instead is the substance: their real strengths, the actual gap, and the next development step, without the box name attached.

Start with the critical roles, the ones whose sudden vacancy would hurt. Use the 9-box grid to find who already combines strong performance with real potential. Then build them a real development plan, not just a name on a list waiting for an opening.

Wujha

Keep up with the ideas shaping better workplaces.