
Endocrinological diagnostics is the set of strategies used to measure and interpret hormonal signals and related markers in order to demonstrate a functional defect or excess, identify its site along the regulatory axis, and quantify its clinical impact in terms of risk and need for treatment. Unlike many other areas of laboratory medicine, in endocrinology the measured analyte is often a dynamic fraction of a complex system: it changes over time, may be bound to plasma proteins, may be peripherally converted into active or inactive metabolites, and may be influenced by drugs, comorbidities and physiological conditions such as pregnancy, age and puberty. For this reason, the “quality” of a result depends not only on the analytical method, but on the entire chain extending from the clinical indication to sample collection, from sample transport to measurement, and finally to translation of the result into a clinical decision.
A highest-level approach starts from one principle: the laboratory is not an oracle, but an amplifier of clinical reasoning. The right test can confirm a strong hypothesis, quantify severity, guide treatment and monitor its safety. The wrong test, or the right test performed under uncontrolled conditions, may generate false positives, false negatives and, above all, “fragile” diagnoses, meaning diagnoses that are not reproducible. This page describes the general principles governing test selection, control of preanalytical conditions, understanding of methods and clinical interpretation. Detailed protocols for dynamic testing, biological variability and analytical interferences, and endocrinological imaging are developed in dedicated pages.
The first step in endocrine diagnostics is not ordering a panel, but defining the objective precisely. In practice, every laboratory request should answer a question belonging to one of four categories: confirming or excluding a suspected diagnosis, localizing the site of dysfunction along an axis, assessing severity and risk of complications, and monitoring response to and safety of therapy. These categories require different tests and, above all, different interpretive thresholds. A “borderline” value may be sufficient for monitoring therapy, but insufficient for an initial diagnosis; a screening test must prioritize sensitivity, whereas a confirmatory test must prioritize specificity. Without clarifying the objective, even a technically perfect result becomes clinically ambiguous.
The second step is to estimate the pre-test probability, because hormones are among the analytes most exposed to the problem of prevalence: many conditions are rare, symptoms are often nonspecific, and reference ranges inevitably include a proportion of “abnormality” in healthy individuals. When tests are applied in a setting of low pre-test probability, an epidemic of false positives is almost inevitable, with a cascade of investigations and iatrogenic anxiety. Conversely, in a setting of high pre-test probability, a value within the range does not necessarily exclude disease if the test was performed at the wrong time, with an inadequate method or under physiological conditions that alter the analyte. Excellent diagnostics therefore integrates two pieces of information: the strength of the clinical phenotype and the reliability of the measured result.
In endocrinology, time is part of the test. Many hormones follow circadian and ultradian rhythms or are secreted in a pulsatile manner. Interpretation therefore requires the request to specify when the sample was collected and under which conditions: fasting or postprandial state, posture, acute stress, sleep, recent physical activity, phase of the menstrual cycle, pregnancy, fever or acute illness. Without this metainformation, numbers lose their physiological anchoring and become difficult to compare with thresholds and operational criteria from guidelines.
The preanalytical phase includes patient preparation, sample collection, type of tube, transport times, centrifugation, separation of serum or plasma, storage temperature and number of freeze-thaw cycles. For some endocrine analytes, particularly labile peptides, catecholamines and some steroid precursors, preanalytical errors are a common cause of uninterpretable results. An excellent laboratory may measure a degraded analyte precisely, but precision does not correct the biological error introduced before measurement.
Control of posture and stress is often decisive for hormones affected by sympathetic activation or by changes in plasma volume. The choice between serum and plasma, and the presence of specific anticoagulants, may also influence some methods. In addition, the timing of centrifugation and separation from the cellular component is crucial when the analyte may be metabolized or adsorbed by cells. For this reason, in clinical settings in which the result changes important decisions, the laboratory request should be accompanied by operational instructions, and the service should have standardized procedures for the most sensitive tests.
Management of the patient receiving therapy is another critical point. Many substances alter the real concentration of the hormone, its free fraction or its immunological measurement. Exogenous glucocorticoids may suppress the hypothalamic-pituitary-adrenal axis and make function tests uninterpretable; estrogens increase binding proteins and may modify the ratio between total and free concentrations for some hormones; drugs that induce or inhibit hepatic enzymes change clearance and steady-state concentrations; high-dose biotin may distort numerous immunoassays. In highest-level diagnostics, ongoing therapy is not a detail: it is a variable that can reverse the meaning of the result and must therefore always be collected and communicated.
Clinical endocrinology has historically relied on immunoassays because they allow automation, high throughput and sustainable costs. However, immunoassays measure “immunological recognition” more than chemical identity, and this makes them vulnerable to cross-reactivity, interfering antibodies and matrix variations. These vulnerabilities become clinically relevant when analytes are measured at low concentrations, when the patient has endogenous antibodies, or when the molecule exists in multiple forms, with similar metabolites or isoforms. The practical consequence is that two different methods may produce non-overlapping results, even if both are technically valid.
Liquid chromatography coupled with mass spectrometry has become a reference for many steroids and for analytes in which specificity and accuracy are decisive. Its main advantage is greater structural specificity, which reduces cross-reactivity and interferences typical of immunoassays. However, mass spectrometry also requires expertise, rigorous controls, standardization, and may be influenced by preanalytical factors, extraction, ion suppression and the choice of internal standards. In addition, it is not always available within clinically useful time frames, and harmonized clinical cut-offs do not always exist on a large scale. Excellence in endocrine diagnostics does not mean indiscriminately replacing immunoassays, but knowing when immunological precision is sufficient and when chemical identification is necessary.
Another crucial issue is harmonization. In endocrinology, interpretation is based on thresholds and operational criteria often derived from studies conducted with specific methods. If a laboratory changes method or platform, cut-offs may not transfer automatically. For this reason, centers aiming for the highest quality must ensure metrological traceability when available, participation in external quality assessment programs and internal procedures for method comparison, especially for analytes that guide irreversible decisions such as surgery, chronic replacement therapy or chemotherapy.
In endocrinology, it is essential to distinguish between reference intervals and decision thresholds. The reference interval describes the distribution in a selected population; it does not automatically define health or disease for an individual. Decision thresholds, instead, are limits derived from clinical outcomes, risk of complications or diagnostic performance of a test under specific conditions. Confusing these two concepts generates frequent errors: a value “within range” may be pathological in a context of altered feedback or in a suppressed axis, whereas a value “outside range” may be physiological in pregnancy, puberty or in the presence of binding variations.
Reference intervals must be specific for age, sex, puberty, pregnancy and, whenever possible, analytical method. The lack of adequate intervals in some populations is a relevant cause of misclassification. In addition, the presence of non-Gaussian distributions and the existence of individual functional ranges, often narrower than the population range, mean that longitudinal comparison within the same patient, under standardized conditions, may be more informative than cross-sectional comparison with a laboratory limit.
A world-class approach uses three interpretive levels. The first is physiological compatibility, meaning consistency with rhythms and sampling conditions. The second is axis coherence, meaning the relationship between trophic hormone and peripheral hormone and, when useful, between total and free fractions. The third is clinical significance, meaning impact on symptoms, target organs and future risk. This triple filter helps reduce overdiagnosis, especially in subclinical abnormalities, and increase sensitivity in cases in which a single number may be misleading.
A robust diagnostic strategy proceeds in phases. The screening phase uses high-sensitivity tests or clinical rules that identify who deserves further evaluation. The confirmation phase uses more specific tests, often repeated under controlled conditions, and dynamic tests when necessary. The localization phase determines whether the problem is primary to the peripheral gland or central, and in some cases identifies a focus of autonomous secretion. The etiological phase distinguishes between autoimmune, genetic, iatrogenic, neoplastic, infiltrative and functional causes. Each phase requires different tools, and a typical error is jumping directly to imaging or genetics without having solidly confirmed the biochemical dysfunction.
In axis-based reasoning, the pair formed by trophic hormone and peripheral hormone is the basis, but it must be interpreted carefully. A “non-suppressed” TSH with high FT4 does not automatically mean central hyperthyroidism or thyroid hormone resistance: it may reflect analytical interferences, drugs, binding abnormalities or non-thyroidal illness. Similarly, a “low” cortisol in a non-standardized sample does not diagnose adrenal insufficiency without considering time of day, stress, binding proteins and measurement methods. For this reason, localization and etiological definition must rest on reproducible data and on selected complementary measurements, not on extended panels.
Dynamic tests have a central role when secretion is pulsatile, when the axis has reserves that must be challenged, or when a basal value is poorly informative. However, a dynamic test is a clinical procedure as well as a laboratory procedure: it requires eligibility criteria, preparation, protocol standardization and interpretation based on method-specific cut-offs. There is no single universal cut-off valid for every laboratory and platform, and interpretation must consider context, pre-test probability and the effects of drugs and comorbidities. The general principle is that a dynamic test should be performed only when the response will change the clinical decision and when the conditions make the result interpretable.
In endocrinology, incongruent results are common and often arise from immunometric interferences, heterophile antibodies, anti-hormone autoantibodies, macrocomplexes, biotin interference, anti-reagent antibodies or matrix effects. The most important clinical feature is not the absolute value, but the discordance between phenotype and laboratory result or between correlated analytes. When a number does not “behave” as expected, the priority is to suspect a measurement problem before concluding that a rare syndrome exists. This approach avoids improbable diagnoses and inappropriate therapies.
Practical strategies include repetition with an alternative method, dilution and recovery studies, use of blocking reagents for interfering antibodies, precipitation for macrocomplexes when appropriate, and assessment with reference methods such as mass spectrometry in selected analytes. Collaboration with the laboratory is an integral part of the pathway, because many interferences are not visible to the clinician without structured dialogue. Knowledge of the typical interferences affecting individual analytes is also fundamental: some tests are particularly vulnerable, and for some clinical scenarios the choice of method is already a measure to prevent error.
A cross-cutting issue is biotin. The use of high-dose biotin, through supplements or specific therapies, may produce falsely elevated or falsely reduced results in numerous immunoassays that use biotin-streptavidin interactions, with patterns that may mimic hyperthyroidism, hypopituitarism, gonadotropin abnormalities or cardiac marker alterations. Prevention requires targeted history-taking and, when necessary, temporary discontinuation according to laboratory indications and guidelines, or the use of methods not susceptible to the interference. The key is to transform a potential risk into a standard procedure: ask, document and manage.
Highest-quality endocrine diagnostics requires an ecosystem: clinician, laboratory and information systems. At laboratory level, internal controls, external quality assessment, management of deviations and change-control procedures when a method changes are fundamental. At clinical level, it is essential to reduce unnecessary variability: same time and sampling conditions for follow-up, same platform whenever possible for longitudinal comparisons, and accurate documentation of therapies and physiological conditions. At system level, reports are needed that include consistent units, intervals appropriate for the method and interpretive notes when known risks of interference exist.
Continuity of care in endocrinology also depends on the ability to compare measurements over time. When a patient changes laboratory, platform or method, systematic differences may appear and simulate worsening or improvement. An excellent system minimizes these untracked transitions and, when they are unavoidable, makes them explicit. This is particularly important in chronic therapeutic pathways and in conditions in which titration depends on small but clinically significant variations. The end result is diagnostics that produces not only numbers, but reduces errors, optimizes resources and improves outcomes.
Finally, modern endocrine diagnostics must be conceived as a process of continuous improvement. The emergence of new platforms, new therapeutic antibodies, new interferents and new indications requires constant updating and dialogue between clinicians and laboratory specialists. Quality is not a static attribute, but a property of the system that is maintained through procedures, audits, training and a culture of error. This approach is what makes it possible to be reliable in simple cases and, above all, in complex cases.