Every mid-sized or large company in Brazil operates, by legal obligation, an occupational health program. Regulatory Standard 7 requires the Occupational Health Medical Control Program, with admission, periodic, return-to-work, job-change, and dismissal exams, each resulting in an Occupational Health Certificate. A company with thousands of employees generates, over the years, tens of thousands of these exams, performed at a variety of different providers spread across the cities where the company operates.
The result is a paradox. The company and the occupational medicine service accumulate an enormous volume of health data about their worker population, and at the same time can't see that population. They have thousands of certificates and no usable population view, neither for risk management nor for prevention.
The data that exists and doesn't become information
The occupational health program produces clinical data at scale. Each periodic exam generates laboratory results, measures, and evaluations that, in principle, describe the health state of the workforce over time. This set would have direct value for management: identifying risk trends in the population, anticipating health problems related to the activity, directing prevention programs to where they matter most, measuring whether the interventions work.
In practice, this value almost never materializes, because the data arrives fragmented. The exams are done at different providers, each with its nomenclature, its units, and its reference ranges, and the result typically returns to the company as a certificate, a document attesting fitness, and not as structured data. What the company accumulates is a pile of certificates, not a population database. The information that would describe the workforce's health exists, dispersed inside those documents, but it isn't in a form that allows seeing it together.
Bates and colleagues, in analyzing the use of data in healthcare, point out that the value of clinical data for population risk management depends on it being structured and aggregable, and that data remaining in unstructured formats doesn't sustain that kind of analysis.¹ Occupational health is an exemplary case: the data exists in volume, but the form in which it exists prevents it from becoming population information.
The cost of not seeing the population
The inability to see one's own worker population has concrete costs, even if invisible in the immediate budget.
From a risk management standpoint, the company that can't aggregate its occupational health data operates blind to the real risks of its workforce. Trends that would only appear in population analysis, a rising prevalence of a certain risk marker in a specific job function, for example, remain invisible, and the company loses the opportunity to intervene before the risk materializes into leave, illness, or liability.
From a prevention standpoint, corporate health programs depend on knowing where to concentrate effort. Without a population view, prevention becomes generic, applied uniformly for lack of data that would allow directing it. The company invests in occupational health fulfilling the legal obligation, but doesn't reap the preventive return the data it itself generates could offer.
From the worker's own standpoint, fragmentation means their occupational history doesn't accompany them in a useful way. Each periodic exam starts from scratch, without structured comparison to the previous ones, and the health trajectory that could signal a work-related problem early is lost between providers and documents.
Why volume doesn't solve it alone
There is an intuitive assumption that, given enough data volume, the population view would emerge naturally. Occupational health disproves this assumption. The volume of certificates is enormous and grows with each exam cycle, and still the population view doesn't emerge, because volume isn't the same as comparability. Tens of thousands of exams in heterogeneous formats don't add up to an analyzable base; they add up to a bigger pile of incomparable documents.
What's missing isn't data. It is the layer that makes the data comparable to each other: that recognizes the same marker measured at different providers is the same marker, normalizes the units, documents the reference ranges, and organizes everything into a structure that allows seeing the population. Vest and Gamm are explicit in showing that the integration and aggregation of clinical data across distinct sources depend on semantic structuring, and that without it fragmentation persists regardless of volume.² Occupational health accumulates volume without accumulating comparability, and that is why it can't see.
The digitalization context reinforces the opportunity. eSocial and the digitalization of occupational health and safety increased the pressure for structured data in occupational health, creating the incentive and the initial infrastructure to handle this data in an organized way. But fulfilling the obligation to report isn't the same as being able to see the population: reporting requires submission, while seeing requires comparability. The company can be up to date with submission and still blind to its own workforce.
From the pile of certificates to the view of the workforce
The change this scenario requires is to stop treating the occupational exam as a document to be filed and start treating it as data to be structured. When each occupational exam result is harmonized, with its marker identified, its unit normalized, and its reference range documented, the pile of certificates transforms into a population base. The company comes to see trends, direct prevention, and measure interventions over data that describes its workforce for real.
This is a case in which the harmonization layer makes more sense as infrastructure than as internal construction. A company or an occupational medicine service is unlikely to internally develop the ability to interpret reports from hundreds of different providers, map markers to recognized standards, and normalize units. This is a data infrastructure problem, resolved by a specialized layer that transforms the dispersed exams into a comparable base.
This is where OpenHealth Technologies operates. The platform automatically correlates multiple data streams with rigorously validated logical layers of laboratory tests, transforming exam results from different providers, in any format, into structured and comparable data, mapped to LOINC, across over 3,500 biomarkers. For occupational health programs and corporate health healthtechs, this means the pile of dispersed exams transforms into a usable population view, able to sustain risk management, directed prevention, and measurement of interventions over the workforce.
Learn how your organization can transform tens of thousands of dispersed occupational exams into a usable population view, the foundation for risk management and directed prevention.

