Earlier this year Zwift introduced Category Enforcement. In races with category enforcement turned on, you can only enter your own category or higher – and this category is determined by your Critical Power, which is calculated by looking at any efforts you have done in the last 60 days between 2-50 minutes long.
Learn all about Category Enforcement here >
Category Enforcement has generally been a huge success. The majority of Zwift races now have it enabled, and the historical problems of sandbagging are now really a thing of the past. However, of course there is always scope for improvement! Here are the key problems with Zwift racing prior to the CE rollout, to have a look at what areas still need some work.
| The Problem | Solved by CE | Solved by suggestions below |
| Riders can enter lower categories than they should | ✔️ | |
| ZwiftPower only looks at 20-minute efforts, so avoiding 20-minute efforts can keep your category lower than your ability | ✔️ | |
| ZwiftPower needs 3 x 20m efforts, so only doing 2 can keep your category low | ✔️ | |
| Strong short duration power is a key Zwift asset, so certain rider types can benefit from high 3-5 minute power | ✔️ | |
| Only efforts in races are considered | ✔️ | |
| Races are normally won by sprints, but sprint ability is not factored in | – | – |
| Categories are determined by w/kg, when raw watts is often a better predictor of success | ✔️ | |
| By splitting on w/kg, the heaviest riders at the top of each category (below A) will be the strongest | ✔️ | |
| There is a need to introduce a ‘minimum wattage’ so super lightweight riders are not overly penalised | ✔️ | |
| Typically riders have very consistent race experiences – always at the front, always in the middle, always at the back | ✔️ | |
| The only real predictor of success is success itself – results should ultimately determine ranking and therefore race pen | – | – |
| Riders need to understand why they are placed in a certain pen | ✔️ |
As you can see, Category Enforcement has delivered some real improvements. By taking a range of efforts in across different durations, and considering all Zwift activities, it is now very difficult to fool the system – either deliberately or simply due to a lack of races.
However, there are many issues still to overcome. Long term we know that Zwift are working on a results-based system, but I understand that is still in the design phase so could be some time away. Hopefully that is a ranking system that updates based on how you perform against those around you.
In the meantime, there are some quick wins that could elevate the Category Enforcement system to the next level.
Zwift Racer Score
I recently came across this research paper which uses a formula to create a ‘compound score’ for a rider’s capability. This score is, in effect, a combination of w/kg and raw watts, to better predict performance versus either w/kg or raw watts alone.
Zwift could use the same formula to create a ‘Racer Score’ which would then be used to determine your category. This removes a number of issues related to communicating w/kg, and better groups riders on ability as demonstrated by the paper. The Critical Power calculation introduced by Category Enforcement would still be the key input to this formula, which is actually very simple:

Example calculation for a rider with a CP of 310w and a weight of 73kg:
Zwift Racer Score = 310 x (310/73) = 1316
Customisable Pen Boundaries
The Racer Score approach certainly improves matters, but ultimately with a fixed category system, race experiences will be very consistent. That’s great if you are always at the front of the field, but for a bottom B or bottom C the experience is really not enjoyable or motivating.
By allowing race organisers to determine how the pens are split (choosing upper Racer Scores for each pen) sometime riders will be at the front of a field and sometimes at the back. This approach is also critical for a future ranking system to ensure variety to keep rankings fair and accurate.
Communicate Racer Score
The calculated Racer Score should be communicated on both ZwiftPower and your Zwift profile. A rider needs to know the result of the calculation so they can understand why they are placed in a certain pen. Simply show the racer their Score! If a racer is inclined to work out their calculated Critical Power, of course they can do this by reversing the formula using their weight in kg.
This will also avoid any confusion with the existing ZwiftPower categories, which still use w/kg boundaries. Medium-term there would be no real reason to keep the Zwift Power categories.
That’s it! 3 simple (in my mind at least!) changes that would massively improve category enforcement.
Questions or Comments?
Share your thoughts below!
Fantastic suggestions James.
OK – not the panacea of results based categorisation – but I think this would make great improvements. And they all look easy to implement suggestions, I hope Zwift open to listening. 🙏🏻
Funnily enough in the DIRT Racing Series I use a very closely related algorithm (but only against 20m watts and wkg) to try and balance the teams across different leagues and it’s proved to work quite well.
First, I wouldn’t use a study that is trying to establish a threshold value in this way as a way of getting a meaningful ‘score’.
Second, I would take a closer look at page 2 of the linked paper and the correlation values found and plots used. Showing, basically, your point about the predictive benefit of raw watts vs W/kg.
Third if I was for some reason absolutely convinced a new score system is a better predictor for category enforcement use on Zwift (which in my opinion this paper doesn’t begin to show) I would formulate it as watts divided by square root of mass which seems, among other things, simpler to intuit/understand — as simply reducing the importance of rider mass vs the traditional W/kg estimation.
I like it. I would guess even people that are consistently at the top of our categories would prefer something like this. I don’t want to have to think about my category or how an effort will affect it, but it’s a reality because that category will determine your racing experience for months. A ranking system like what you described will encourage everyone to continuously improve since it’s essentially ONE category without concrete category boundaries.
The ability to have custom/narrower ranges within a cat is a great idea. TFCs Mad Mondays take an upper/lower cat approach. Yes, it relies on people not bagging up front but organizers also DQ people. Best fun I’ve had in a race in a long while.
Two things…..
I agree 100%. I can go over 3,9w/kg, very rarely and only if I am at peak race fitness, have a very good day and nearly kill myself ….. as long as the race is short. There is no way I could do that for an hour or do it in shorter races consistently. I think taking top performances over a three month or so period, is far better than a one off race of any distance. I think that average should be based on the performance at that race distance too. Yes, riders can back off to not get upgraded but for me that ruins the experience. It is bad enough being stuck in a category you cannot even stays with riders but for real B riders it is harder as the difference between the fastest and slowest is much greater. Maybe there should be more cats or promotion and relegation based on top 10% and bottom 10%.p, say every month
I’ve said this before and I’ll say it again. The only time W/kg is useful a useful metric for determining rider speed is on a hill with a slope of more than 7-8%. At any other time it’s Watts vs CdA that determines your speed. Zwift knows your Watts, Zwift knows your CdA (weight and height both affect this) therefore Zwift knows what speed you are capable of.
For some reason Zwifters have become obsessed with W/kg as the be all and end all of performance metrics when in reality it is only appropriate on steep climbs and can only provide a rough indication at all other times.
W/kg, Zwift Racer Score and absolute Watts limits for light riders are unnecessary if you calculate and determine the categories based on the actual speed someone can ride at in Zwift rather than waste time trying to invent some other formula which will only provide an approximation.
Zwift knows the formula for speed. Zwift should use that formula. There is no need to invent any other formula.
And then Zwift should junk it all for a proper results based system because performance based categories are stupid.
Or they should just categorize us based on our results in races. Sure, they’d need to do something with new people so there isn’t a constant stream of new people who should be A’s winning races in lower divisions) until they earn enough points to race as A’s (maybe cat enforcement as currently done for the first 5-10 races until you have enough results for accurate placement?), but it would solve all of this and be very visible and easy to understand rather than Zwift either using their hidden algorithm and no one knowing how they got their score or publishing their algorithm for speed determination (which they won’t want to do as they consider it a trade secret).
I think the category system makes racing boring and worthless.
I believe that without a category system all problems would cease to exist.
This is the easiest answer.
We’ve had this discussion in other forums, but you do realize that you’re free to race in A or uncategorized races and not pay attention to any of this, right?
On another note, Herd Winter Racing is uncategorized, 7 race times on the weekends, probably starting right after the first ZRL season.
As someone who just got bumped up to B (and is at the very bottom of B) and has been struggling in the Herd Summer Racing Series, I’ve been eagerly waiting for the Herd Winter Racing to begin.
One of my favorite series to race because I have no expectations of winning, so there’s no pressure. Just try to hold onto the front group as long as I can, and, as soon as I get dropped, I try to hang with the next fastest group until they drop me. Every week, it’s my goal to last a little longer. Perfect series to get used to racing at the bottom of a faster category while I try to get stronger.
I was sort of hoping that the new season would start after the current Herd Summer series ends in 3 weeks.
I don’t agree with the sentiment that it’s all boring and worthless, but it’s arguably worth highlighting that people have -very- different understandings of what ‘fair’ is. Zwift is probably never going to be fair and probably should focus on fun.
It would be technically trivial to do a geolocation routine and figure out altitudes and temperatures to estimate what kind of condition riders are riding under. You could even probably estimate something like median income, and make a good guess of what kind of non-sensored gear/nutrition/conditions/advantage they have.
But I don’t think it’s something they are likely to do.
Great article. I was trying to get around CE and explain the benefits of it to a group of “ZP Categories” traditionalists beyon md tl the anti-sandbagging. Your table would have helped a lot!
I think one important point you are making for any system is an effort of communication / transparency.
Also CE is a little too volatile as it stands now for competitions happening through a long period of time.
In any case hopefully little by little the case is cracked. Not really because of sandbagging, but because being able to race agains other people from similar strength is what makes most races fun
A weight categorization like boxing?
Results-based Ranking
It’s by far the best way, fixes all the problems with Watt floor and W/kg ceiling issues. Eliminates sandbaggers and cruisers (assuming it’s enforced). Takes care of the issues with some people having high W/kg but no sprint so never being able to compete at the finish, etc. If it is enforced and event organizers can set custom boundaries (or choose to have field split into evenly sized pens shortly before start – so you don’t know what pen you’ll be in in advance), it might fix the issue of people always being in the same position in a race. I’m sure it would have flaws (and ways to be exploited – and we know which teams will find/take advantage of them), but it’s better than looking at one specific timepoint or even a range of timepoints that doesn’t include the power/time that’s critical in most non-hill climb races.
Hopefully, it’s coming.
Cruisers??? You mean not willing to participate w Sandbaggers pushing really hard at the start, then holding promotable watt average for 17.5-18 minutes before backing off enough to not hit Cat limits? A C race shouldn’t be full bore 3.1-3.3 wkg for 40-45 minute race. Bs shouldn’t be 4.1-4.3 either. Cat enforcement eliminated that horse hockey, but then they allowed Cs to hit 4.2 wkg for 5 minutes, you aren’t a C at that level of relative power. Sorry but if you don’t have sprint power you are going to win many races in Zwift or real life unless you race racers well below your fitness level.
I still see people cruising right below the limit for the whole race and then sprinting. They just have to be more careful about it because going too far over once (though the new extra 5% gives them a bit of a buffer) will get them promoted.
With results-based rankings, if they win too often, they get promoted. You either race to get better (and get promoted) or you just take it easy and stay where you are, but don’t win anything.
I wasn’t saying the people with no sprint should magically win, but, if they use results-based ranking, someone with good 20 minute power but no 5-60s power won’t get promoted as fast. That’s not me, by the way, I’m a heavy diesel engine who can do ok on flat sprints, but I see a lot of people complaining about having good distance power, but not being able to stick with the cat they get promoted to because the first surge drops them off the back.
Results based ranking would also take care of the people with A level power racing in C’s because they are lightweight and don’t have the raw W. Results-based will eventually get them to end up where they can compete at the right level. For some of them, that’s high B (because they can either pull away or rest on the hills while others are going all out, so they have the energy to work harder on the flats), for others, it’s a bit lower than that. But, you won’t have someone with 4.3 W/kg demolishing the C field because their FTP is 198 W.
Results-based completely gets ride of the need for raw W floors or W/kg ceilings for any cat because, after they win enough times to get appropriately placed, they’ll be racing with the appropriate groups.
If Zwift HQ makes it super customizable, a promoter could have lots of small pens where everyone is with other similarly competitive people, or they could still have the 4 big pens, but there would be better groupings because if you win too often you move up and, if your results don’t let you compete in that level, the results that put you there eventually expire and you fall back down.
If Zwift were to make it super customizable I think a good place to start would be more climb finishes. In fact, by allowing any race to have an arbitrary start and end point (say the escalator in London) a lot of the concerns about favouring one particular ‘ability’/hack go away. There’s downsides too of course in that people will feel like they don’t know the route as well but I suspect that would be quickly mitigated by the community.
I think that, using custom lengths, promoters can make courses end wherever they want now. Is that not correct? Custom starts might be a challenge with the need for pens and a line to cross that starts the timer, but it might be doable. As a heavy rider who has honed my descending (because I so often get dropped on the climbs), I’d love to have a race that starts at the top of a hill and encourages you to get the most out of the descent.
Question: Would you consider a hybrid ranking system of results based and the one proposed here? My concern with 100% results based ranking is it could provide the wrong incentive of pushing people into races that only favour their strengths (ex. Tiny Races for those with great sprints).
That’s a fair point. I’m the sort of person who races whatever, wherever (I’m a low B who is 103 kg, and every weekend I do a hill climb race even though that’s clearly something I’m never going to win), but I get that different people choose differently (if winning – something I’m never going to do – is their ultimate goal, that makes sense). If someone only races their strengths, they’ll definitely get promoted far beyond their capabilities in their weaknesses and that could lead to siloing where different types of people never interact. Or, it could lead to specialization and people with different, more-defined roles on teams in a team-based racing series, which could be interesting. I don’t know. I’d definitely think you’d need something like Critical Power to put people in an initial spot until they have enough races for results-based ranking to take over so that you don’t have a constant stream of people who should be A’s starting out as Ds until they win enough to be promoted and making D’s unbearable for the actual Ds.
Being a light rider (55kg), the w/kg alghoritm used by Zwift is far distant from real effort required on real roads. Even with large groups, the draft effect is much lower than reality.
The thing is, there will be and still are people who work around the category enforcement rule to stay in a lower category to win. Until there is a enforcement to move those people up to the next category there isn’t going to be any real change…
Results based ranking with enforced minimum categories.
Did my first CE race today. Much, much, much better! Although not perfect it’s a vast improvement.
I just hope that the new season of WTRL incorporates it – no three strikes and you are out either- go over the limit and you get DSQ and moved up.
It is odd that CE is the new thing Zwift is pushing and the flagship racing series doesn’t seem to use it. Maybe it’s just that right now there is no place where your cat is being advertised and the boundaries are changing.
I had a weird race in March and then again in May that bumped me up to A. It wasn’t easy, or always fun, but racing with the As was a real learning experience. First, they either ignore or chuckle at the ZP guys. Second, I learned a lot about how to be a smarter, more efficient racer. This morning some B prima donnas were getting after a fast but legit B, calling him a cheater, etc. it was pathetic. I think guys should stop obsessing about the few remaining cheaters and work on their weaknesses. My sprint has come a long way since March. I think my punch has gone from about 20% to over 70%. Just work harder boys! Having said that, if Zwift introduces any new features that interfere with my slow climb up B cat I will be right there with the prima donnas crying foul!
hhy
I don’t understand why people think a results based ranking will be a pancea. In Zwift, where there is no penaly nor cost for entering a race, results based systems have significant flaws. There’s two options as I see it; all races count towards ranking or just best results. The flaws/loopholes in the first are obvious; racers looking for a lower category can just ride what would normally be group or pace partner rides as races, finished lowly and downgrade; stupid. If the latter, then this – effectively – punishes riders who race more as freakish result are more likely. My highest ever race rating for an event is currently when I raced up in an A cat Crit. 5 riders, I came dead last (I was trying to improve my 20 min power, the rating was a complete surprise). A donkey on a bike could have achieved that. But even if this is resolved by a better scoring system, why would a company want to make the game hardest for those who use it the most?
I agree, and disagree. I also think there are some significant issues with a results based system, but I believe they can be addressed through design. It’s certainly not as easy as some are making out – other games have been trying to land on the perfect system for years.
Something like iRacing of Microsoft TrueSkill seems to be the best approach.
Shh, something that doesn’t exist is perfect and will definitely solve all our problems. Don’t fall for the trap of just enjoying the good things that are happening to you right now in the present; keep thinking that only in the future will things be perfect!
> The majority of Zwift races now have it enabled,
Not Quite – ZwiftHacks says 622 races in next 7 days. Only 268 of those with CE. Before we move on to ticking some of the other boxes, lets just get all the races converted first.
Interesting, I certainly thought it was more than that. I wonder if there are some particular race ‘types’ that are limiting things, as of course it doesn’t fit certain event types.
zwiftinsider article posted a few days before yours:
Back in April 2022, 23% of races were using category enforcement. Today, that number has doubled to 46%. (This is based on the next 7 days of events as of the date of this post. Currently there are 705 races happening in the next 7 days, with 327 of those races using the new category enforcement tools.)
Whilst possibly an improvement, this won’t stop those idiots that have cheating at the forefront of their minds. I am mid-CatB according to ZwiftPower, but CE has decided I’m an A, no big problem since I’m mostly doing TTs against myself. In a CE event yesterday I was squarely gubbed by a Cat B who averaged 6.7W/kg, and seemingly weighed 32kg. CE doesn’t appear to have quite got them in the correct category.
Developing a model that correctly models all aspects of bike racing, indoors or outdoors is impossible.
On the other hand, basing categories on success (or lack of) has been used for better or worse IRL since bike racing began. It is not perfect, but it is straightforward to explain and it is hard to complain when you are moving up because you are doing well and placing high.
We can quibble over how to do the ranking, but at least basing it on results should result in fewer complaints.
Lots of good ideas here. Great you are talking about them James. I think that additional categories are the answer A-; B+; C+ etc a bit like what DIRT series and FRR uses. Smaller cats provide more competitive racing experiences.
I also think nothing stated so far accounts for what races people like to do. As a triathlete, i enjoy racing iTT’s – a sure fire way to up your watts, vs those who like draft pack racing with sprints. But even in series events, its never the winners who seem to Cat up, which can’t be right. That just wouldnt happen i football for example! So some race ranking score that isnt just volume driven seems very useful.