Establishing and Enabling a Center of Production Excellence

By: Nick Travaglini

Software is in a crisis. This is nothing new. Complex distributed systems are perpetually in a state far from equilibrium, operating in what Richard Cook has called a “degraded mode.” It’s through a combination of technical artifacts, organizational practices and policies, and pure gumption that they manage to maintain themselves through time.
However, there are some organizations that seem to have an easier time of it than others. Resilience is an activity they perform, rather than a property attributable to the organization. They don’t take that achievement for granted.
Software is both a product of and part of a sociotechnical system. Unfortunately, the “socio” bit tends to be underappreciated. In this guide, we’ll talk about the concept of a “Center of Excellence” or “Observability Guild” or just a plain ol’ bunch of people who have questions for their system, and how such a group can bolster the greater organization and help it achieve production excellence.
Change from within
There are at least two more or less formal institutions that an organization can create to coordinate their observability efforts. The more formal of these is a “Center of Excellence,” while the less institutionalized version is sometimes called a “Community of Practice” or “Guild.” For our purposes, we’ll define them as a functional subsystem within an organization meant to adjust that organization’s behavior, typically to improve some designated dimensions. In other words, the goal is to understand an organization from the inside so well that the group can engage in constructive criticism. For the sake of concision, we’ll refer to this as a Center of Production Excellence (CoPE).
In this sense, they need to have a certain degree of authority and autonomy. This is analogous to safety departments in organizations that attempt to do resilience engineering. According to David Woods, those departments must be:
- Independent
- Involved
- Informed
- Informative
The people who participate must have direct operating experience and come from different parts of the organization so that they can cross-check each other when evaluating the current work processes and preparing interventions and recommendations.
Yet, just because this group has that experience and operates for the sake of improving things, doesn’t mean that rocking the boat might not lead to trouble. Some parts of the organization may understand them as a blessing and others a curse. As such, the group needs access to independent funding and other guarantees that grant it sufficient power to do its work in order to sustain itself—in case certain factions try to impede it.
To build upon this analogy to safety departments, Provan et al. propose that the CoPE should work to induce resilience by “creating foresight about the changing shape of risk, and facilitating action” proactively. They call this “guided adaptability.” What does this look like? We’ll get into that below, but the outcome should be fewer customer-facing incidents for a similar amount of work because they’re preempted. Since those incidents are theoretically not occurring, we can’t count them except maybe as near misses. Besides those near misses, other telltale signs may include qualitative changes in what count as incidents and better response when we do have an incident. We’ll want to track these things to evaluate the efficacy of the institution, since they’ll provide a better signal about how the organization is doing than a lossy metric like the number of incidents.
Differences that make a difference
When it comes to making changes in an organization, Hébert-Dufresne et al. found that it takes both bottom-up and top-down practices. Bottom-up practices are those which are taken up and transmitted horizontally through a network while top-down practices are interventions which introduce an accelerant or a dampener. The bottom-up spread is the main driver of the process, and the top-down inflection makes it easier or harder for the spread to occur; both aspects must be understood and evaluated reciprocally. Furthermore, the researchers note that success will result in a marked qualitative change in the organization: it should be clear once the phase shift has happened.
Start from the bottom
Let’s assume that an organization has already decided to start using Honeycomb, the CoPE has authority to act on its mandate, and someone wants to make a change. Using the model from above, here are some activities that a Center or Guild can do to start and spread good observability practices with Honeycomb and what they may need from the greater organization.
The first thing to do is find as many of those intrinsically-motivated individuals and bring them together in a regular meeting to talk about the CoPE’s mandate, what they want to achieve, and why. This initial group forms the basis of a sociogram, which can be enriched with information about things like periodic sprint cycles, regular post-incident meetings, widely-adopted standards and tools, and other things that don’t get the grease because they’re not squeaking.
Collecting this information up front is crucial because it will inform the development of a twinned strategy involving both passive and active tactics.
Passive tactics
Passive tactics hitch onto already-existing habits and motions. They inflect and latch on, letting another motive force propel them while slowly transforming it with each cycle of repetition. Ideally this is symbiotic, and the best chance of achieving that is to gather sufficient information about how the existing motion is detrimental to the organization’s production excellence goals and introducing Honeycomb to address that deficiency.
All of those tactics are attempts to proactively establish what Laura Maguire calls common ground and to lower the cost of coordination between people. That happens when people study what’s happening in production and communicate regularly about it, and have the time in a low-pressure, psychologically safe situation to ask questions and grow familiar with each other’s knowledge, proclivities, and working styles.
The members of a CoPE should therefore learn about these organizational patterns and use their influence to change them ever so slightly—and reinforce that change until it has become a part of the routine. Once those passive maneuvers are up and running, the CoPE members merely need to perform regular check-ins, or play their own part if they are participants of the relevant social institution. That light touch frees them to sprinkle in more targeted interventions.
Active tactics
The set of active tactics which a CoPE could perform are much more generic and targeted. These include running regular training and enablement sessions like Lunch & Learns, publishing a newsletter, advising on or collaborating on custom instrumentation libraries, and supporting the general observability toolset.
More broadly, since the role of a CoPE is to support the performance of resilience by the organization, this is a good space to consider designs for incident response. This can span from:
- Creating on-call rotations and associated supporting materials, like handover communication methods
- Running an ops-focused office hours
- Building out a Learning from Incidents program
The increasing number of irreducible dimensions within software systems and the changing relations between them requires contributors to adopt an attitude of humility about what they know or may expect to happen. It therefore makes sense to build in mechanisms to foster continual learning and to create an environment where people can and will stretch to fill the gaps that inevitably open.
Part of how the organization can act as a support in this case is to grant the CoPE power to make these changes.
Top to bottom
As for the top-down approach, Hébert-Dufresne et al. are not specific about what behaviors would constitute ones which promote or hinder the spread of the changes that the CoPE makes. In the case of rolling out a product like Honeycomb, we can venture a few ideas backed by research.
One thing that an organization’s management can do is to build the CoPE to be as autonomous as possible. This includes independent funding sources and protection from capricious actions by those whose interests may conflict with the changes that the CoPE is pushing. Related to this is a deference to frontline expertise. As Woods says, this expertise can be expressed as initiative and observed as people deviating from established plans or routines. In a case where deviation does occur, it’s important for management to validate it as appropriate and to have safeguards in place that prevent blame and punishment.
Returning once again to the safety analogy from above, a common shorthand for this is letting anyone halt production if something ‘on the ground’ seems dangerous. In software, this can manifest in ways like empowering anyone to declare an incident or deferring/dropping planned work in favor of things with longer-term benefits, like instrumentation or conducting incident retrospectives. Management must create buffers which support the CoPE if it deems these activities necessary.
Another way that management can contribute is to look at unusual backgrounds or skills as assets, not distractions. Scott Page has famously demonstrated in books like The Diversity Bonus that teams with diverse experiences and “lower” individual aptitude perform better than homogenous ones where each individual is “better.” As such, organizations should seek internal and external candidates for roles who don’t pattern-match well with the group as currently constituted.
Finally, people in management should consider voluntarily leaving their current roles. Copious research, like studies conducted by Julian E. Orr and Ruthanne Huising, have found that knowledge silos form between hierarchical levels in organizations, not just between departments or teams. This is a major problem for organizations because ossified power dynamics inhibit good communication, those who wield power for too long may use it to protect themselves instead of for the good of the organization, and those people may personally suffer debilitating effects like falling behind in their own technical skills. To counter this, one might consider adopting the model of the Engineer/Manager Pendulum or other techniques of rotating leaders like sortition.
It’s all about adaptability
A CoPE can be a powerful means for an organization to initiate transformations which foster resilience as it matures and its environment changes. In order to do this, its design, activities, and supporting structures require careful consideration. An organization’s agents must have a true desire to change in order to make appropriate decisions in those regards, and they must empower it to do the work of guiding adaptability.
The precondition for guiding adaptability is adaptability. This is the capacity or potential for agents to modify their behaviors, mental models, and priorities as the complex, dynamic system that they participate in changes. It occurs when agents re-evaluate their circumstances and, determining that the current state of affairs is insufficient to achieve their goals, draw upon heretofore novel resources which change what they can affect. A CoPE helps increase that capacity by making those resources available and preparing agents to use them when needed.
Organizations can aid in this process by adopting policies which amplify that work and avoid dampening it. In doing so, they demonstrate and enact their commitment to production excellence and should expect to reap the benefits.
Independent, involved, informed, and informative: The characteristics of a CoPE
As Honeycomb’s Field CTO Liz Fong-Jones says, production excellence is important for cloud-native software organizations because it ensures a safe, reliable, and sustainable system for an organization’s customers and employees. A CoPE helps organizations cultivate the practices and tools necessary to achieve that consistently.
We’ve already analogized the CoPE with safety departments. David Woods says that those safety departments must be:
- Independent
- Involved
- Informed
- Informative
Here we’ll elaborate on what each of those characteristics means, why the CoPE should also match those qualifications, and how to achieve that status.
Why it matters: Power dynamics
We’ll begin by discussing the quality of independence. This will allow us to think about how the CoPE relates to other parts of the organization and why that matters. It’ll also help us define some criteria for determining how independent your CoPE is at any given time.
What do we mean by “independence”? Independence is the capacity to autonomously perceive or affect something. There can be no independence if the things being compared are absolutely identical or completely different. Independence is established by mediating structures like tools, formal roles, and life experiences.
Independence and organizations. So far we’ve discussed organizations as if they are a coherent entity, and in a certain sense, that’s correct. However, there are undoubtedly tensions within organizations as well. The effect of this is that there is not—and cannot be—perfect “alignment” between all facets of an organization. The people who work there each have some changing degrees of independence. Therefore, the coherence of an organization is a dynamic result—not a static premise to be assumed.
This brings us to why it matters and what it means for the CoPE to be independent.
There are some within the organization who benefit from the current state of affairs and have motives to keep it that way. They may therefore attempt to inhibit the CoPE because it is intended to change the state of affairs and perhaps that will disturb those benefits.
These tensions, however, aren’t all bad. David Stark argues that it is to organizations’ benefit to permit multiple views on what is valuable (and what is the correct thing to do) because dynamism enables adaptation as conditions in the environment change. Without that multiplicity, an organization risks overfitting to the short term and lacking the resources to effectively pivot when it meets a great challenge.
For the CoPE to do its work, it needs insulation from powers which would halt changes it deems necessary. In practice, this may look like funding the CoPE from a customer success or product department budget, rather than engineering.
The CoPE should make sure that its efforts clearly benefit the generous department’s own goals in the long term, that it reports back how it’s working towards those goals, and that this aid entitles the department to participate in the CoPE’s work as a crucial stakeholder.
If we can’t get rid of these tensions, then we should make good use of them. We’ll discuss below how to balance this tension over time and to utilize it to the organization’s benefit.
Effective communication
This section will focus on the three other qualities that an effective CoPE should display: involved, informed, and informative.
Involved. The CoPE should participate in the work processes that they are attempting to change. A good heuristic here is: don’t ask someone to do something that you won’t do yourself.
We want to avoid the CoPE taking an outsider’s perspective on the situation and merely prescribing actions for others to take. In those situations, it’s all too easy for this to turn the CoPE into mere counters or regulators. These are impractical and wasteful uses of scarce resources. Moreover, punitive measures can create distrust between CoPE members and their colleagues, leading to alienation and secrecy.
That sort of culture, as Sidney Dekker has argued, is toxic. That toxicity may prevent the CoPE from achieving its stated objectives by ceasing communication about the conditions of work, thus stopping crucial flows of information about the things it is attempting to change. Backchannels like private Slack channels or DM conversations are anathema to the whole endeavor.
The CoPE should help implement whatever changes it proposes and intimately understand the challenges of the work.
Informed and informative. The CoPE, as something that isn’t identical to the teams and work processes it’s affecting, has something to learn and something to teach. In terms of learning, it needs to study and produce knowledge about actual behaviors (“Work-as-Done” (WaD) to use researchers’ language) of people and the broader system. Teaching comes when the CoPE communicates the actual state of affairs to other stakeholders and affects the organization’s decision-making.
These two correlated aspects make the CoPE a supremely helpful advisor to frontline teams and organizational leadership. Understanding WaD lets the CoPE coordinate between disparate teams, enabling smooth working relationships. This “common ground” makes it easier to make mountains into molehills.
However, this may require slowing or quickening work timelines. A CoPE can only affect those deviations from plans when it is independent from leadership and can speak candidly about the good, bad, and ugly. Leadership has to be willing to hear bad news, or the CoPE needs to be insulated enough that it can deliver that news when it’s unwelcome.
Staffing up your CoPE
Getting the right people working in the CoPE is crucial to success because these change agents must limber up the organization and promote the flexibility necessary to perform resilience.
We’ll look for teammates who share enough in common to work well together, but who don’t necessarily perfectly overlap so that they can play off each other’s strengths. We’ll also want them to come from a place of real commitment to ensuring great operations and a mindset that emphasizes the customer experience.
Finding people who match those characteristics doesn’t need to be hard. Here’s a strategy that’ll help you find and recruit these colleagues.
The challenge
The working members of your CoPE face a daunting challenge: by joining, they agree to the task of modifying a complex, adaptive system which is subject to financial, human, legal, and other constraints that have placed it in a locally optimal position. For the organization, things are going pretty well, all things considered. It’s normal—and expected—for things to turn out alright.
The CoPE members’ job is to convince the organization to change now, proactively, rather than wait until later (when things start going wrong) and thus the change will be a reactive one that takes more effort. That’s a tough ask because:
- It’s not clear that the organization is anywhere close to the point where things aren’t going so well
- When things normally go well, it’s easy to chalk up cases where things go poorly as one-offs
- Until things do go poorly, it’s hard to determine what to do that will preempt a big qualitative change from good to bad
The solution
To meet that challenge, here’s what we recommend:
- Enlist distinguished engineers (or adjacent positions) who already believe that changes in observability are needed.
Making a change requires people who will advocate for that change. It also requires that people listen to those advocates. The easiest way to get that is to find people who are already respected in their parts of the organization and who believe in the desired change.
If your organization is already using Honeycomb, then its History, Enhanced Reporting, and Activity Log features can provide good data to start your search. These three data sources allow you to see what people are querying (and how complex those queries are), how much they’re querying, and what else they’re doing in Honeycomb, respectively. That data can help you triangulate on relative “power users.”
An important qualification is to look for people who will provide what Scott E. Page calls “diversity bonuses” once they are assembled. In The Diversity Bonus, Page lays out an argument that any group addressing complex challenges like those the CoPE will face requires a wide range of backgrounds, heuristics, information, and skills. Each of these added to a group provides a bonus because it changes how the group can engage with the problem space, and they can also be combined to form synthetic methods of working.
The details of the particular problem(s) that the group is addressing matter, so this diversity should be intentionally constructed. Thankfully, the book lays out guidelines about how to produce and accrue the greatest bonuses in various types of problems, like predicting or innovating. Consider making use of this theory when making staffing decisions.
- Once the members are recruited, bring them together in a regular meeting to talk about the CoPE’s mandate, what they want to achieve, and why.
This initial group forms the basis of a social graph. That graph should represent social relations which can be enriched with information about normal engineering practices that each person and their team perform, such as periodic sprint cycles, regular post-incident meetings, or widely adopted standards and tools. It may also include informal relations like regular coffee meetings, carpool groups, and the like.
Let the group talk about their work and trust them to figure out, together, what needs to change and how. That conversation may benefit from a facilitator to ask open-ended questions and keep the discussion from stalling. As an output, the group should author a charter which includes a set of problems that it needs to scope and address, one or several stopping rules for its work, and rationale for these decisions.
Collecting this information up front is crucial because it will inform the development of a twinned strategy involving both passive and active tactics for making change.
- Get sign-off on the cooperatively-drafted charter.
The CoPE may face headwinds from various parts of the organization; hence the need for committed resources and a location that permits it the leeway it needs. This means that clear expectations with its immediate neighbors, like a sponsoring department, are crucial. CoPE members can’t do their work if they don’t have the material, substantive resources necessary to make change.
Each member and the CoPE’s immediate ‘upstream’ supporters have to reach some clarity in order to coordinate. Therefore, a group review and commitment to follow through on the charter is paramount. Then, and only then, is the CoPE ready to get to work with the confidence and power to engage in creative destruction.
The CoPE and other teams, part 1: Introduction and auto-instrumentation
The CoPE is made to affect, meaning change, how things work. The disruption it produces is a feature, not a bug. That disruption pushes things away from a locally optimal, comfortable state that generates diminishing returns. It sets things on a course of exploration to find new terrains which may benefit it more—and for longer.
Laurent Hébert-Dufresne and his co-authors produced a model of social organization. That model shows the transmission of changing behavioral norms, which can result in an overall bolstering of fitness. We take this model as our guide. Based on their research, we believe that only through coordinated and mutually reinforcing changes can an organization successfully achieve its goals. These changes should promote prosocial behaviors which decenter short-sightedly individualistic achievement.
The following sections consist of recommendations for a CoPE. A CoPE may use these to drive towards those desired outcomes; they especially focus on observability and the effective use of the Honeycomb product. The recommendations feature several socio-technical reforms at the level of individuals and teams (the bottom of the organization’s formal hierarchy), and institutional and policy changes at the management level (the top of the formal hierarchy).
We’ll cover “bottom-up” and “top-down” practices, which should mutually reinforce each other. The defining characteristic of bottom-up practices is that they diffuse through an organization’s network of frontline practitioners. Meanwhile, the telltale sign of top-down practices are that they are interventions that introduce an accelerant (or a dampener, as the case may be). Additionally, the CoPE can interact with their colleagues in the lower position in the hierarchy in two different ways: passive and active. Each of these “modes” of engagement should affect the behavior of the team. However, the distinction between the two is how they affect them. The passive mode consists of the team receiving materials produced by the CoPE and incorporating them into their existing habits and dispositions themselves. The active mode consists of the CoPE inciting breaks from existing habits and dispositions.
The long-term impact of these interventions should be “cooperation,” which Hébert-Dufresne et al define as “behavior that carries group-level benefits.”
Telemetry instrumentation
The foundation of good observability is the instrumentation of your software. This instrumentation takes the form of additions to your services’ source code, either as software packages like libraries or new code handwritten by developers.
Our general recommendation is that folks instrument all of the services available to them in order to achieve at least basic visibility into their system’s behavior and performance characteristics. The next step is to customize their instrumentation to emit more detailed and information-dense telemetry. This instrumentation should produce what are, effectively, structured logs in JSON files. Each of these files is the content of what Honeycomb refers to as “events.”
Honeycomb, as an observability tool, allows developers and other interested parties to analyze that telemetry data by means of “queries” and to model their software system. Its columnar datastore and query engine allow its users to write new queries without needing to pre-aggregate data or to index it in advance. Furthermore, Honeycomb’s model encourages users to create what are termed “wide” events that make use of “high cardinality” values.
Events
Wide events are structured logs with lots of dimensions; Honeycomb supports events with up to 2000 dimensions at time of writing. Each dimension consists of a key:value pair. That key is an “attribute” or “field” and is analogous to a column header in a spreadsheet. The more dimensions in a dataset, the more possible ways that one can segment and analyze the data.
Cardinality and dimensions
High cardinality refers to a quality of the data, specifically describing the values in each dimension. High-cardinality data is more detailed than low-cardinality data. What makes something “high” or “low” cardinality is the elements in a set. A set with only three elements has lower cardinality than a set with 100000; the value True in a set of {True, False, null} is less detailed than the value 20109 in a set of {0, 1, 2, ..., 99999}. The latter is more distinctive and distinguished, and therefore, more informative.
For example, consider two events consisting of the same two dimensions. One dimension records whether a user is logged in and its data type is a Boolean, while the other records their user ID and is a string. Now suppose that each user logs in at the same time. The events will only be distinguishable by the user ID value. Building upon this, suppose that instead of two users, there are two thousand and all of them log in at the same time. All of those events are only distinguishable because of the higher cardinality set of possible user IDs. That greater distinguishability is what makes dimensions with higher cardinality data more informative, in the sense that dimensions help us answer more questions, in the form of: “Is user X logged in, yes or no?”
Together, these serve as a basis for comparing the “information density” of different events. Wider events with higher cardinality values are more useful for analyzing system behavior because they enable finer-grained distinctions between segments and more flexibility in the scale of analysis. In other words, one can twist and turn the model in more ways and zoom in and out to a greater extent.
Honeycomb recommends instrumenting your software with OpenTelemetry (OTel) in two ways: the first makes use of OTel’s auto-instrumentation capabilities, and the second builds upon OTel with custom work tailored to your system.
OpenTelemetry auto-instrumentation
The open-source OpenTelemetry project offers a wide array of libraries and a robust suite of other tools which are useful for instrumenting software. Since instrumentation depends on your organization’s services, we’ll only briefly touch on this aspect. Instead, the focus will be on the CoPE’s strategy for growing adoption of the span-and-trace-based mode of observability, distributed tracing—OTel’s sweet spot.
The problem. In our experience, many developers struggle to adjust to tracing. This may be due to experience with metrics or (unstructured) logs—they have trouble breaking from expected design patterns. Or it may be that they haven’t yet switched from prioritizing their system and its components (which is what many tools focus on) to prioritizing their customers’ or users’ experience. Or perhaps they simply don’t grok how to connect the data analysis part to system performance.
The solution. The first way to address this problem is to auto-instrument everything. Every service. Every proxy. Every library. (Ok, some things like ColdFusion monoliths can’t be instrumented. But do the things that can!)
Developers will want access to data that represents what’s happening with the services that they immediately work on or have responsibility for. Without that, there’s no way to help move them to care about production excellence.
Once that data is available in Honeycomb, the next step is getting developers to look at and make use of it. To that end, we recommend holding Introduction to Honeycomb trainings as one active tactic, and to build in structural constraints to existing processes which create the conditions for passive adoption.
The method. The precise constraints are, again, context dependent on the specific patterns of behavior that the CoPE wants to inflect. One example: the pull request “show me.”
Many software development organizations rely on Git-based technologies. A frequent pattern is to use pull requests (PRs) to gate merging new code on a peer review, and part of submitting that PR is completing a templated form explaining the rationale and important details about the proposed code change.
One thing that a CoPE may do is to modify this template to include a section requiring a Honeycomb query. The query should display the before, and running the query again once the code change is merged, should display the after. This will let everyone involved check how the code change has ‘moved the needle’ and affected the customer experience. Think of it as analogous to unit testing.
This benefits adoption because it requires the developers to learn how to represent their system via Honeycomb’s query builder and to focus their attention explicitly on how the production system’s characteristics impact customers. It also documents this for future reference and makes it transparent to all involved.
The CoPE and other teams, part 2: Custom instrumentation and telemetry pipelines
Auto-instrumentation can get teams started, but you can’t rely on it alone. This section discusses its limitations in more detail and how a CoPE can help teams overcome them.
Custom instrumentation
Once your teams begin working with telemetry from auto-instrumentation, they’ll soon realize something: it’s reeeeally barebones. While the specifics vary by language and library, most of the official OTel options are generic. They may provide information in a span about things like http.method, http.status_code, or the name of a unit of work, but that’s a far cry from anything specific to your business case. Auto-instrumentation is good for basic information, and because it includes the pieces needed to construct a trace-like span hierarchy.
However, to get fine-grained details about your system’s behavior, your teams will need custom instrumentation—but improving telemetry in this way is a “wicked” problem, meaning there is no solution that doesn’t leave some problematic remainder (did someone say tech debt?).
The problem. Organizations want high-fidelity data about what their system’s doing because it’s more useful than low-fidelity data. In Honeycomb, that means that they want wide events with lots of dimensions (key-value pairs) and high-cardinality data for values.
Auto-instrumentation provides narrow events with some high-cardinality dimensions and a set of industry-standard semantic conventions. To get better data requires additional work. It takes time for a team to decide on the appropriate units of work, to add the instrumentation code into the main business logic, and to put in the creative effort to come up with the new fields to write in. Furthermore, teams writing their own instrumentation may deviate from one another in the semantics that they use when creating those new fields. These deviations can make it difficult to establish and maintain common ground between individuals and teams.
Finally, it’s impractical to attempt to instrument everything, given the tradeoff between the time devoted to instrument code vs. the return of value. This indeterminacy is similar to the halting problem in the theory of computing: there’s no clear way to know that one has sufficiently instrumented their code. If that’s the case, when conditions change for the business (e.g., new services are added corresponding to new business imperatives), then the instrumentation will need updates.
The solution. In light of these challenges, most organizations only see a mountain of work. However, our guiding model suggests that there is a tipping point where this goes from active effort to passive habit. It’s the role of the CoPE to get their organization to—and through—that point.
The way to go about this is two-fold:
- Make the approach to the summit easier to traverse, and
- Give the hikers the appropriate gear and provisions for the trek.
The method. One of the best ways to address this challenge is with SLOs! Honeycomb’s SLOs can be built upon data already produced from auto-instrumentation, but teams often find that auto-instrumentation doesn’t provide enough insights to really keep the parts of the system as reliable as they’d like. It also may not provide the information that partners like the product team want when using SLOs to make choices about prioritizing new features or reliability work. If enough people agree that it’s not sufficient, the CoPE can step in.
The first step will be to convince the product team (and other parts of management) that the lack of information means that they need to invest time in instrumentation. Reliability isn’t just about whether the team is meeting its SLOs today—it’s about if they’ll meet them in the future, too. Having good enough instrumentation to even make that determination is reliability work.
Once time has been allotted, then comes the fun part. A CoPE can start to address #1 (make the approach to the summit easier to traverse) by defining a set of semantic conventions. Martin Thwaites advises creating some that are flush with the OTel conventions when they abut (e.g., when adding fields related to HTTP traffic, make sure to prefix http.*). From there, the conventions should make sense locally, whether that’s relative to the teams or the larger organization. The CoPE will need to draw upon their own background working within the system and the thoughts of their colleagues to build this out.
With these conventions as axioms, the next step is to determine what exactly should have additional instrumentation, and what context should be included. As noted above, this will change over time. It’s the role of the CoPE to help teams create these and to figure out when they stop being effective.
To help teams start, the CoPE should build a custom library for each of the languages in use through the organization. These will supplement the auto-instrumentation already in place. The initial version will again draw on the CoPE members’ experiences and informal contributions from their colleagues. Later revisions should be driven by the teams that rely upon the libraries. But if no team “owns” the resource, then who will do that work? Won’t it fall prey to the “tragedy of the commons?”
Fortunately, that scenario is a fictionalized, simplified ideal—and in fact is largely an edge case, as Elinor Ostrom argues. The key to both the governance of this common resource—and of prompting teams to improve it—is to create forums where it is discussed as the solution to a problem. This helps frame it as something worth putting effort into. The ideal setting for this is during post-incident reviews.
The type of post-incident review meeting that the CoPE should strive to create is the one advocated by the Learning from Incidents community. In those spaces, incident analysts lead discussions amongst incident responders (and other interested parties) where people explain what they did to respond and how it made sense for them to do those things. It is a chance for “frontline” people to share the expertise earned through normal work in a psychologically safe space, for others to learn about how their colleagues and technical components really work, and to contribute to a discussion regarding where instrumentation was insufficient or outdated.
The CoPE can take this as feedback from the socio-technical system and work with the teams that depend on the libraries to update them, eventually handing this off to the teams completely. Beyond that, and based on what they now know they need, individual teams can add any instrumentation to the code’s nooks and crannies that even the libraries can’t reach.
Telemetry data strategy
An unsung hero in reliability is the pipeline that sends telemetry data to whatever backend analyzes it. While an SRE or platform team will probably take on the bulk of the work managing it, a CoPE can make a few key contributions to this crucial flow of data.
The problem. One of the great challenges that businesses face is managing the telemetry data their systems produce. There are a host of things to take into account like access control, availability and latency, content (e.g., user privacy and data hygiene), volume, and more. Honeycomb helps make sense of the data it receives, but that’s conditional on what actually makes it to the tool.
The heart of the problem is that many organizations don’t recognize that different data has different value to different teams. For example, sending events with lots of PII can be extremely valuable to an organization because it’s high cardinality and thus lets the team scope investigations quite tightly. However, certain regulatory environments encourage finding other means to achieve that scoping, so if the organization begins to do business in such an environment, then the PII’s value drops considerably.
The solution. A CoPE can begin by investigating and creating processes for sharing the value of different data throughout the organization.
Engineering is tasked with managing tradeoffs in the pursuit of a goal. In this case, telemetry data isn’t equally valued by all, so the relevant hierarchies of value must be accounted for in order to understand the ways they conflict. Once those conflicts are known, then the parties involved can negotiate towards the pursuit of their objective(s). Those negotiations produce a data strategy, which can serve to determine suitable tactics and techniques. Thus, a CoPE can facilitate the creation of this strategy and the derivative tactics that the organization will use.
The method. In a certain sense, a CoPE has already begun this work by its very nature. As noted before, drawing staff from across the organization will bring local knowledge about these values into dialogue immediately. So a CoPE can begin developing a data strategy with its own constitutive expertise.
However, it can’t end there. Other people need to be interested in how the organization works and resist the drive to myopic focus on just their own tasks and team. The organization needs incentives for curiosity. Once those are in place, conducting workplace studies into the values of a locality and sharing the results of that research is a fine way to circulate knowledge.
A CoPE’s guide to alert management
Alerts are a perennial topic, and a CoPE will need to engage with them. The bounds of this problem space are formed by two types of alerts:
- Reactive alerts (in Honeycomb, we call these Triggers): they are alerts that fire after some event, like crossing a pre-determined boundary.
- Proactive alerts (Burn Alerts based on Honeycomb’s SLO feature): these give notice before crossing a threshold; in the case of SLOs, that means before failing to meet the stated objective.
Understanding what these alerts are and how to configure them is one thing. Thinking about what they each do for your organization, and how using one or the other affects things, is another.
Evaluating the utility of each type of alert
The great challenge of alerts is how to get them to the right people at the right time. That’s because an alert conveys information signals, but if it isn’t transmitted to the right person at the right time, then it’s just noise.
Such a situation is, in fact, much more prevalent with Triggers than with Burn Alerts. Why?
Triggers are extremely particular in the signal that they convey. Consider the following scenario: you’ve set up an alert to sound if a load-bearing column in a building exceeds its safe capacity. In such a case, it’s totally appropriate for someone who receives the alert to respond “So what?” Only people with prior knowledge about the situation can answer that question. To them, it’s information. To everyone else, it’s noise.
Burn Alerts generated by SLOs, by contrast, are much more likely to prove informative to a general audience. That’s because context is built into them through the SLO. A Burn Alert effectively tells its receivers that they either have some amount of time before too many bad things will have happened, or that an unusually large number of bad things have happened in a certain window of time and that’s putting them at risk of having too many. Returning to the column example above, it’s like getting a warning as the load is getting to be too heavy.
Burn Alerts and SLOs inform their receivers about what the organization values and the hierarchy of those values. They let organization members know what to prioritize when load shedding is necessary. For example, it’s simple to decide between working on a new feature or stabilizing API latency when a Burn Alert has indicated that you have 12 hours until you miss your API’s SLO for the month. This creates a shared understanding and common ground for communication and collaboration.
This is why Honeycomb advocates so strongly for using SLOs. That shared understanding works perfectly with a well-defined north star: customer experience. Working together to map critical user journeys and then establishing landmarks via SLOs with proactive alerts is the paradigmatic setup for achieving production excellence.
The problem. Despite these differences, many organizations don’t effectively differentiate their alerts. Reactive alerts are almost certainly the most common type used, and their volume and lack of context hinders prioritization. This induces alert fatigue and harmful stress.
A CoPE must understand that alerting is crucial to their success, so they need to develop and implement a solid alerting strategy.
The solution. A CoPE should endeavor to make all alerts into Burn Alerts; in practice, this isn’t really possible, so the goal should be to optimize the ratio of Burn Alerts to Triggers given the organization’s needs.
Every Trigger is an operational risk. They indicate that the organization hasn’t created a strong enough structure that depersonalizes necessary information and activity, meaning that the organization has a bus factor. They also indicate a relative concentration of power, because these are the people who benefit from information asymmetry. The ones who can answer that “So what?” question for a given Trigger are a pocket or silo within the organization—but our model accounts for that, and acknowledges them as necessary factors in complex systems. With that in mind, the best thing to do is to make them explicit and learn how to work them to the organization’s advantage.
The method. The first thing for a CoPE to do is to convert as many Triggers to Burn Alerts as possible (don’t forget to take into account the difference between Exhaustion Time alerts and Budget Rate alerts). This means reformulating them into forward-looking goals, like the load-bearing column example from above. The new alerts should then route to public spaces, like a general #Ops channel in Slack or dedicated channels for specific on-call rotations.
Any remaining Triggers should be routed to private DMs for the one or small set of people who have the relevant context. This makes sure they get relevant reactive alerts without creating alert fatigue for others. These Triggers should be reviewed regularly with their recipients to check if they’re still necessary or if they can be made into SLOs or even just deleted.
Finally, the CoPE needs to build in social institutions that manage the pockets of asymmetrical information. We recommend the Learning from Incident-style incident reviews because their whole point is to induce the circulation of information. When people feel that they can speak openly about what they did and how it made sense to them in a way that benefits everyone, they become encouraged to share their power and thereby redistribute it throughout the group. In this case, that power is information, and sharing it makes for a more resilient organization.
A CoPE’s duty: Indexing on prod
Odds are that a software engineer today is really focused on one place: pre-prod. Short for “pre-production,” this is slang for an environment where software code operates in a prototype phase of its development lifecycle.
Common sense would have one believe that this is a safe space, a workbench of sorts, where problems can be found and remediated. Then, once engineers are reasonably certain everything’s working properly, they advance it to a matching environment called production, where the code behaves like it did in pre-prod and it merely needs to be managed by an operations team.
That story is a comforting lie.
The problem. A wise woman once said: “I always test my code and then I test it in production, too.” The truth is that code in prod and code in another environment may look the same, but behave differently. Prod and “lower environments” are purposely different because of the social aspects of their existence as sociotechnical systems.
Consider load testing. In a pre-prod environment, the traffic used to test performance is generated and runs along code paths that are executed using synthetic data created by an organization’s employees. In prod, the traffic is generated by actual customers and users. Now sometimes, the employees can also use prod—but most of the people using it won’t overlap with that category. So they’re mostly different people who aren’t reducible in their knowledge, mental models, interests, etc. to the employees. What they do with the code will also be different and will change over time, so they’ll affect the system in different ways.
The tests run in lower environments will be different from what actually happens in prod. This means teams need to shift their focus to prod and build up their experience and tooling to effectively support it.
This contravenes much of the history of software engineering. So many of the different methodologies and technologies treat user behavior as disruptive and destructive. The underlying assumption is that the goal of engineering is to produce a system that works, and that’s so difficult to achieve that hermetically-sealed workspaces are required. But if the working system is so brittle that it can’t handle what customers and users do with it, then it’s as good as worthless to the organization. Organizations don’t just set a normative standard of behavior for their customers and users. They also serve customers and users. In other words, they dynamically co-create each other.
The solution. The way forward here is to treat prod as a sense organ for the organization, like a person’s ears or nose. It’s a medium for receiving information that the organization can process and then turn into definite actions that engineers and product are particularly attuned to. A CoPE therefore needs to ensure that these groups have the right signals coming into Honeycomb and that they’re transformed and shared with the rest of the org.
Some of this will look like what we’ve discussed before in terms of instrumentation and alerting. However, there are distinct interventions that a CoPE should consider if they find that the organization implicitly values work in lower environments and that it’s necessary to shift their colleagues’ center of gravity to prod.
The method. First things first: developers need to understand what effects their code deploys and releases have. Once they have that info, then they can circulate it amongst their team, the eng org at large, and finally, find ways to share it with the broader organization. The key to succeeding is to start small and let iterative changes compound.
Honeymarkers
One track that a CoPE might take is to begin adding Honeymarkers automatically as code deploys. These appear as annotations on visualizations in the Honeycomb UI and can include information like the deploy ID or commit ID. They’ll allow anyone to correlate behavioral changes with deploys, making it easy to see if a code change had the desired effects or if it needs to get rolled back.
Custom instrumentation
Then, a CoPE could help teams to add that same ID as an attribute via custom instrumentation so they can actually include it in their query parameters. That permits triangulating between the ID, the marker, and any other parameter(s) serving as the dependent variable(s).
Permalinks
Once those are in place, it becomes very easy for a team to put that PR “show me” that was suggested earlier. The Honeycomb feature that makes this work and serves as a documentary papertrail is the URL permalink, which allows anyone to revisit query results indefinitely.
Sharing out to the wider org
Finally, engineers can begin sharing their changes in places like sprint reviews, departmental all-hands, and even company-wide events like demo days.
These practices all reinforce the idea that engineering is working on things that improve the experience of the system’s users and that align with the organization’s goals. They also break down knowledge silos, because explaining what a change is and why it was done requires cross-functional contextualization (i.e., storytelling). Furthermore, they give everyone a chance to celebrate the excellent work they’re doing and express appreciation for one another’s achievements.
Additional steps beyond this point might include utilizing feature flags to distinguish between deployed and released code, or a mechanism like GitHub Actions Deployment Protection Rules to gate changes on the results of Honeycomb queries.
Determining a CoPE’s efficacy—and everything after
A Center of Production Excellence (CoPE) is a more or less formal, provisional subsystem within an organization. Its purpose is to act from within to change that organization so that it’s more capable of achieving production excellence. We’ve focused mainly on how best to construct such a subsystem and what activities it should pursue. To conclude, let’s return to the point of a CoPE, discuss signs of success, and evaluate the impacts it’s having.
Signs of success
Returning to the model developed by Hébert-Dufresne et al., the authors note that neither a bottom-up first nor top-down first approach is sufficient to address the greatest challenges, which require coordinated group change. Each end of the spectrum must act mutually and reciprocally, and allow themselves to change for—and with—the other. As they make their incremental changes, there will come a point where a dramatic qualitative change occurs in the holistic system.
That qualitative change is increased cooperation through pro-social or pro-group behaviors. The actors involved are understood to require top-down institutional support to achieve this because those pro-group behaviors provide only indirect benefits. To sustain, compound them, and produce habitual cooperation requires not just their own activity, but also a conducive environment. As such, a CoPE should aim to grow pro-social behaviors amongst the organization’s members regarding practices that assist and inform their colleagues in their work to diagnose, improve, and maintain system performance. Exemplary activities include:
- Adding instrumentation according to established conventions.
- Revising those conventions once they’ve gone stale.
- Building dashboards to help on-call engineers quickly jump into action or onboard newcomers.
- Ensuring deploy markers are accurately denoting deploys and teams are checking the effects of their code changes once released.
These alone are not enough. In order to build towards that qualitative shift, the organization will need to adopt specific policies so that people repeat these things and establish a trend. Accolades and rewards may need to be rethought. For example, the NBA tracks and rewards players when they assist a teammate in scoring. Only when both happen in an amplifying cycle can a CoPE turn their colleagues’ behaviors into habits and make the desired change.
Measurements, or emotional support numbers
By now, it’s well-established that measures are corrupted once they’re put to work (cf. Goodhart’s Law). That said, people still seek them out and they do perform a significant psychological function: promoting a feeling of agency and control in low-trust environments.
Organizations may insist on tracking something, so it’s good to understand the transformation that a CoPE is making so that we can create measurements around its impact.
Let’s note right away: the number and duration of incidents, unqualified by any other attributes, is a bad way to go. Much research has shown that categories like “incident” are constructed and change as sociotechnical factors shift. This makes univocally determining what an incident is and when it started/stopped nearly impossible, so we should reconsider the idea that the best way to track this change is with an extensive measure.
What a CoPE is really after is making things better (a qualitative change). Is there a number or some other method of indicating that? Indeed there is!
Think of this change as an intensive one. Intensive properties, like density or temperature, are ratios of extensive properties. That makes them dependent on the context in which they’re situated. As the extensive quantities change, they induce changes in the intensive property, and vice versa; this change is analogous to the categorical changes mentioned above.
Given the number of factors at play in an organization, producing a single number is a significant challenge. Instead, organizations can look to dynamic modeling techniques—similar to how meteorologists model weather systems—to track several vectors and show their relationship via a heatmap. This provides a high-level view of a system’s qualitative changes, which can then be analyzed to produce visualizations of more particular relations. Hébert-Dufresne et al. take this tack in their own research, using a heatmap that, as they say, “highlights the phenomenon of institutional localization in which a given institutional level dominates the fitness landscape in some subset of parameter space.”
The switch to thinking in terms of intensive transformations may be outside the norm, especially if management is used to receiving One Big Number™. However, the adoption of techniques like this facilitates clear communication regarding the actually-desired information. That’s in contrast to using lossy signals like number of incidents and MTTR. Those are poor proxies and it doesn’t do anyone any good to report on them simply because they’re familiar.
Nonlinear progress
A final word on what it takes for a CoPE to succeed, addressed directly to organizational management: heed these words.
Your position as an overseer in the formal hierarchy grants you certain powers. You often have the power to hire and fire, and own a budget. Don’t mistake that as sufficient. David Woods has called organizations “tangled layered networks.” This means that your formal position and power is only one layer among many others that compose your organization; you are yourself multi-functional and operate across layers simultaneously. Each layer has its own topology, and therefore, its own lines of force.
If you have decided that a CoPE is right for you, then for it to succeed, you may need to make certain tradeoffs against your position in the formal hierarchy in favor of another layer in the network. Don’t be afraid to do this. A true sign of leadership is knowing when to defer and to follow others. Be open to seeing the changes wrought by the CoPE as solutions to problems rather than as a problem to be solved, and to changing your own ways of working and being in your organization. This is part of why the Engineer/Manager Pendulum—and techniques like sortition—can be worth considering. Swapping positions in one network may serve to address problems and allow for nonlinear progress.
Conclusion
A CoPE is meant to solve problems. That means that it is a productive tangent, branching off from the limit that had been reached heretofore.
A tangent isn’t a linear continuation. If the organization could continue on linearly, then it wouldn’t be facing a problem. But it is, so something different is required. With the right approach, and the right partner, a CoPE can be the solution to your organization’s production excellence problems.