In many areas of modern life, people depend on indicators to understand situations that are too complex to observe directly. A company may measure productivity by counting how many tasks employees complete. A university may use graduation rates to evaluate its academic performance. A hospital may record waiting times to determine how efficiently patients receive care. In each case, a number represents something larger: productivity, educational success, or service quality. Such indicators can be useful because they make comparison possible and can reveal patterns that would otherwise remain unclear.
However, a problem can appear when an indicator stops being only a description and becomes a target. Once people know that their success will be judged according to a particular number, they have a reason to change their behavior in ways that improve that number. Sometimes this produces exactly the desired result. At other times, the number improves while the underlying reality changes very little. The indicator may then become less reliable precisely because people are trying to perform well according to it.
Consider a customer service department that measures employees according to the average length of each telephone call. Managers may introduce this measure for a reasonable purpose. If customers are waiting too long, shorter calls could indicate that employees are solving problems more efficiently. At first, the measure might provide useful information. Employees who understand common problems and communicate clearly may indeed complete calls more quickly.
The situation changes, however, if employees are rewarded mainly for keeping calls short. An employee facing a complicated problem now has two goals that may conflict. One is to help the customer completely; the other is to finish the call quickly enough to maintain a good performance score. Some employees may begin ending conversations before every issue has been resolved. Others may transfer difficult customers to another department. Average call length could fall, suggesting greater efficiency, while customers actually need more calls to solve the same problems.
This does not necessarily mean that employees are dishonest or that managers have chosen a useless measure. Before call length becomes an important target, it mainly records behavior. After rewards or penalties are attached to it, it can also shape behavior. The measurement is therefore no longer simply observing the system from the outside. It has become one of the factors influencing what happens inside it.
Similar effects can occur even when nobody deliberately tries to manipulate a measure. Imagine a school that wants to improve students' writing. Teachers introduce regular writing assessments and use the results to identify areas that require attention. Initially, the assessments may help teachers understand students' strengths and weaknesses. But suppose the school's reputation later depends heavily on scores from one particular writing test. Teachers naturally have a strong reason to prepare students for that test.
More classroom time may then be devoted to the types of essays, topics, and structures that commonly appear in the assessment. Students' scores may rise considerably. Yet interpreting this improvement is not simple. Perhaps students have genuinely become better writers and can transfer their skills to unfamiliar situations. Alternatively, they may have become especially skilled at producing the kind of writing rewarded by the test. Both changes can produce higher scores, although they do not mean exactly the same thing.
The distinction matters because an indicator normally captures only part of the phenomenon it represents. Writing ability, for example, includes organization, clarity, vocabulary, revision, and the ability to adapt to different purposes and audiences. Any single assessment must emphasize some of these elements more than others. As attention becomes concentrated on what the assessment rewards, aspects that receive less attention may remain unchanged or even become less important in daily practice.
Organizations sometimes respond by adding more indicators. If one measure gives an incomplete picture, several measures may seem more reliable. A workplace might evaluate employees according to speed, customer satisfaction, accuracy, and cooperation instead of speed alone. This can make it harder to improve one number by ignoring everything else. Yet a larger collection of measures brings its own difficulties. Employees must decide how to divide their attention, and managers must determine how much importance to give each result. A system designed to capture more of reality can therefore become more complicated without necessarily becoming more informative.
There is also a less visible consequence of measurement. What an organization chooses to record can influence what people consider worthy of attention. Activities that appear regularly in reports, rankings, or performance reviews are difficult to ignore. Other activities may be valuable but harder to express numerically. Their effects might develop slowly, depend on context, or become visible only after several years. When immediate, measurable results receive most of the attention, such activities can gradually move to the margins.
None of this makes measurement unnecessary. The absence of indicators creates problems of its own. Without them, organizations may rely too heavily on impressions, personal preferences, or memorable individual cases. A numerical measure can reveal patterns that are difficult to see otherwise. It can show whether performance has changed over time, identify differences between groups or departments, and give people a common basis for discussion. The difficulty lies in deciding what conclusions a number can reasonably support.
For example, a hospital may succeed in reducing the amount of time patients wait before receiving attention. That change may represent a genuine improvement. But waiting time alone cannot show whether patients received appropriate treatment, whether staff experienced unsustainable pressure, or whether another stage of the process became slower. These possibilities do not make the original measure meaningless. Instead, they change the questions that decision-makers need to ask about it.
Time also matters. A measure can be highly informative when it is first introduced and less informative later. People learn what is being evaluated and adapt their routines accordingly. In some cases, this adaptation is desirable: the measure encourages behavior that the organization wanted from the beginning. In others, people discover ways of improving the visible result without producing an equivalent improvement elsewhere. Often, the difference between these situations becomes clear only by examining consequences beyond the indicator itself.
Measurement therefore creates a continuing tension. Organizations need simplified representations because they cannot pay equal attention to every aspect of a complex reality. Yet those representations influence decisions, and decisions influence behavior. As behavior changes, the meaning of the original measure can change with it.
A number should consequently be understood in relation to the environment that produced it. A rising score, a shorter waiting time, or a higher rate of completed tasks may all contain valuable information. What each result means depends partly on what people did in response to being measured and on what remained outside the measurement. The challenge is not simply to obtain better numbers, but to understand when better numbers still represent the broader improvement they were intended to capture.
# Reading Comprehension Test
### Instructions
Read the text **“When Measures Become Targets”** carefully. Then answer questions 1–15.
For each question, choose the option that is **best supported by the text**. Select only **one answer** for each question.
### 1. Which statement best expresses the problem described near the beginning of the text?
**A.** Indicators become unreliable whenever they reduce a complex situation to a number.
**B.** An indicator can become less representative when people change their behavior in order to improve it.
**C.** Organizations usually select indicators that measure a different phenomenon from the one they intend to evaluate.
**D.** People tend to reject indicators once those indicators begin to affect their performance.
### 2. What does the author mean by saying that measurement has become “one of the factors influencing what happens inside” a system?
**A.** Organizations eventually modify their goals to match the indicators they have selected.
**B.** Employees become responsible for determining how their own performance should be measured.
**C.** The act of measuring can affect the behavior that is being measured.
**D.** A measure becomes useful only after consequences are attached to the results.
### 3. A delivery company evaluates drivers mainly by the number of packages delivered each day. Drivers begin leaving packages in unsafe places because returning them to the warehouse would lower their totals. How would the text most likely interpret this situation?
**A.** The indicator has encouraged behavior that improves the measured result without necessarily improving the broader service.
**B.** The indicator has become useless because delivery totals have no meaningful connection with successful delivery.
**C.** The drivers' response shows that personal judgment would provide a more reliable measure of delivery performance.
**D.** The higher totals should be treated as evidence of improved efficiency unless customer satisfaction also declines.
### 4. Why does the author distinguish between higher writing-test scores and broader improvement in writing ability?
**A.** To suggest that assessments should evaluate every component of writing equally.
**B.** To show that students who practice a particular form of writing cannot transfer their skills to unfamiliar tasks.
**C.** To argue that a single assessment cannot provide useful evidence about students' abilities.
**D.** To show that improvement on a measure may reflect either broader development or adaptation to what the measure rewards.
### 5. A research center evaluates its staff using publication speed, accuracy, collaboration, and the usefulness of completed projects rather than publication speed alone. Which interpretation is most consistent with the text?
**A.** Using several measures prevents any one aspect of performance from influencing researchers' behavior.
**B.** Using several measures can represent more dimensions of performance, although researchers may still adapt their behavior to what the system rewards.
**C.** Adding measures makes the evaluation more reliable because each measure corrects the limitations of the others.
**D.** Adding measures transfers responsibility for defining successful performance from managers to researchers.
### 6. A factory tracks defective products to identify production problems. Later, employees begin avoiding unusually difficult products because these are more likely to increase their defect rates. Which interpretation best applies the text's reasoning?
**A.** The measure has revealed that difficult products should not be included when evaluating normal production performance.
**B.** The measure remains equally informative because the defect rate accurately records the products employees actually make.
**C.** The workers' response demonstrates that the original measure was poorly designed because useful indicators should not affect behavior.
**D.** The measure may now represent production quality differently because employees have changed what they produce in response to being evaluated.
### 7. What is the relationship between the customer-service example and the school example?
**A.** Together they show that a measure can influence behavior in different settings even without requiring dishonesty.
**B.** The first shows harmful adaptation, while the second shows that adaptation to a measure normally improves the broader outcome.
**C.** The school example limits the customer-service example by showing that only workplace incentives create conflicts between a measure and a broader goal.
**D.** Both examples suggest that the measures were unsuitable from the beginning because they captured none of the qualities they were intended to represent.
### 8. A university evaluates professors mainly by the number of courses they teach. Professors consequently devote less time to mentoring students, an activity the university values but does not include in the measure. Which explanation best integrates the relevant ideas from the text?
**A.** Professors are choosing an immediate measurable result because mentoring produces benefits only after several years.
**B.** The university has made teaching appear more important than mentoring even though both activities could be measured equally well.
**C.** The evaluation system directs attention toward a visible outcome while leaving another valued activity outside what receives formal recognition.
**D.** The problem results mainly from relying on one measure and would disappear if mentoring were added as a second indicator.
### 9. Why might adding more indicators fail to solve the problems associated with measurement?
**A.** Additional indicators reduce managers' ability to identify which employees have performed well.
**B.** Several indicators usually measure overlapping aspects of performance and therefore provide little additional information.
**C.** People may have to respond to several incentives at once, while decision-makers must still determine how much importance to give each result.
**D.** Once several indicators are introduced, employees can no longer identify which organizational goals should receive priority.
### 10. What function does the paragraph beginning “None of this makes measurement unnecessary” serve in the development of the text?
**A.** It introduces an exception by suggesting that the earlier problems occur only when organizations choose inappropriate indicators.
**B.** It replaces the earlier criticism of measurement with the claim that numerical evidence is generally more reliable than other evidence.
**C.** It summarizes the disadvantages discussed earlier before proposing that organizations increase the number of indicators they use.
**D.** It qualifies the preceding discussion by explaining why measurement can remain useful despite the limitations already identified.
### 11. Consider the customer-service, writing-assessment, and hospital examples together. Which interpretation is best supported by all three?
**A.** An improvement in an indicator can be meaningful while still leaving important aspects of the broader phenomenon uncertain.
**B.** An indicator provides useful evidence only until people begin changing their behavior in response to it.
**C.** Improvements in measured performance often occur because problems are transferred to parts of a system that are not being evaluated.
**D.** Indicators become more informative when they measure outcomes rather than the processes used to produce them.
### 12. Imagine that an organization reports strong annual results on its main indicators, while employees have gradually reduced time spent on useful activities whose effects are difficult to quantify. Which interpretation is best supported by the text?
**A.** The reported improvement is probably inaccurate because indicators cannot represent activities with long-term effects.
**B.** The organization should give the unmeasured activities numerical indicators before deciding whether overall performance has improved.
**C.** The measured results may reflect genuine progress, but they are insufficient by themselves to determine whether the organization's broader performance improved.
**D.** The unmeasured activities should receive less attention because their contribution cannot be demonstrated through the organization's reports.
### 13. Which interpretation best captures the role of adaptation throughout the text?
**A.** Adaptation makes an indicator less informative whenever people become aware of how their performance is being evaluated.
**B.** Adaptation can either support the purpose of a measure or weaken its connection with the broader goal, depending on the behavior it encourages.
**C.** Adaptation is mainly problematic when people intentionally improve an indicator without improving their actual performance.
**D.** Adaptation is useful when it improves the measured outcome, although additional evidence may be needed to determine the size of the improvement.
### 14. An organization wants to determine whether a large improvement in one of its performance indicators represents genuine progress. Which approach is most consistent with the reasoning of the text?
**A.** Examine how people changed their behavior in response to the measure and whether relevant outcomes outside the indicator changed as well.
**B.** Give the indicator less weight and combine it with additional measures before interpreting any improvement.
**C.** Determine whether the indicator still measures the same activities that it measured when it was originally introduced.
**D.** Compare the result with previous years and treat a sustained improvement as stronger evidence of genuine progress.
### 15. Which conclusion is best supported by the text as a whole?
**A.** Indicators are most trustworthy before people understand how the results will be used, because behavior has not yet been influenced by measurement.
**B.** Organizations should interpret measurable improvements cautiously unless several independent indicators show similar changes.
**C.** The main limitation of measurement is that complex goals cannot be represented adequately through numerical information.
**D.** The value of an indicator depends not only on what it measures but also on how its use affects behavior and its relationship with the broader goal.
[[Respuestas del examen de comprensión - Inglés]]