New reparent metrics for Prometheus
We've added two Prometheus metrics that show when a shard's primary changed, and why:
planetscale_vitess_planned_reparents_totalcounts planned reparents, such as the ones that happen during maintenance or when you request one.planetscale_vitess_emergency_reparents_totalcounts emergency reparents, where a replica was promoted to recover from a failed primary.
Both count successes and failures, split by the planetscale_reparent_result label. The planetscale_component label tells you which component ran the reparent: vtorc means VTOrc recovered the shard on its own, and vtctld means the reparent was requested.
Three things to know about the VTOrc metrics:
planetscale_vtorc_recovery_typeholds the name of the recovery VTOrc ran, such asRecoverDeadPrimary,ElectNewPrimary, orFixReplica. Our docs previously listed its values asplannedandunplanned, so update any queries that filter on those.- On Vitess 23 and later,
planetscale_vtorc_failed_recoveries_totalandplanetscale_vtorc_successful_recoveries_totalcarryplanetscale_keyspaceandplanetscale_shardlabels. planetscale_vtorc_detected_problemsreports the problems VTOrc currently sees in your cluster. Join it against the recovery counters to see what VTOrc did about each one.