Skip to main content

Analytical Guide

How Analytical Methods Are Validated

What validation demonstrates: accuracy, precision, specificity, range and robustness, how they relate, and why validated does not mean every batch was tested.

[Research Use Only] This guide is for controlled laboratory research workflows only. It is not for human or veterinary use and does not provide applied-use guidance.

Validation answers one question

Every number on a certificate arrives by a procedure, and every procedure is capable of producing a number whether or not it is capable of producing a correct one. Validation is the work that separates those two cases. It is not a quality mark awarded to a laboratory, and it is not a property the material inherits — it is evidence that a specific procedure, applied to a specific kind of sample, for a specific purpose, does what it is being relied upon to do.

Why the intended purpose comes first

The purpose determines which characteristics matter, and the guideline is explicit that this varies. An identity test and a quantitative impurity test are validated differently because they are relied upon differently: one has to distinguish a substance from other substances, the other has to put a defensible number on a small quantity. Q2(R2) covers common uses including assay, potency, purity, impurity testing as either a quantitative or a limit test, and identity. Asking whether a procedure is validated, without saying for what, is an incomplete question.

  • An identity test must discriminate; it need not quantify
  • A limit test must decide against a threshold; it need not report a value
  • A quantitative impurity test must be accurate and precise at low levels
  • An assay must be accurate and precise across its working range

The performance characteristics

Q2(R2) uses the term performance characteristic — a technology-independent description of a characteristic that ensures the quality of the measured result — and notes that earlier versions of the guideline called these validation characteristics. Each answers a different question about the procedure, and each has a matching limit.

CharacteristicThe question it answersWhat it does not establish
Specificity / selectivityCan the procedure measure the analyte in the presence of what else is there?That the amount reported is correct
AccuracyHow close is the result to the accepted true value?That repeat measurements will agree
PrecisionHow closely do repeated measurements agree with each other?That any of them is near the true value
RangeBetween which values does the procedure perform acceptably?Anything about behaviour outside that interval
LinearityAre results proportional to the true values across the range?That the procedure is accurate at any given point
Detection limitWhat is the lowest amount that can be detected?That a detected amount can be quantified
Quantitation limitWhat is the lowest amount that can be measured with suitable precision and accuracy?That anything below it is absent
RobustnessDoes the procedure survive small deliberate changes to its parameters?That it will survive changes nobody tested
The right-hand column is the part usually left out. Every characteristic is bounded, and validation is the process of establishing where those bounds sit rather than removing them.

Accuracy and precision are independent

This is the relationship most often collapsed, and collapsing it is how a consistent result gets mistaken for a correct one. Precision is the closeness of agreement between a series of measurements of the same homogeneous sample under prescribed conditions, expressed as variance, standard deviation or coefficient of variation. Accuracy is closeness to the accepted true value. A procedure with a systematic bias can be extremely precise and consistently wrong, and nothing in the scatter of its results will reveal that — which is exactly why both are validated, by different experiments.

  1. Precise, not accurateResults cluster tightly around the wrong value. Repeating the measurement increases confidence in an incorrect answer.
  2. Accurate, not preciseResults centre on the true value but scatter widely. Any single result may be far from it.
  3. NeitherScattered and biased. Usually visible, and the least dangerous of the three because it does not look convincing.
  4. BothResults cluster tightly around the true value. This is what validation sets out to demonstrate, and it takes separate evidence for each half.
The first state is the one that causes harm, because it produces exactly the pattern people read as reliability: the same answer every time.

How accuracy is demonstrated

Accuracy cannot be established by measuring something of unknown composition, because there is nothing to compare the result against. Q2(R2) describes approaches that both work by supplying a known: applying the procedure to an analyte of known purity such as a reference material, a well characterised impurity or a related substance, and comparing measured against theoretically expected results; or a spiking study, where a known amount of the analyte is added to a matrix containing all other components and the unspiked and spiked results are compared.

  • Comparison against material of known purity, such as a reference material
  • Spiking a known amount into a matrix of the remaining components
  • Assessed across an appropriate number of determinations and concentration levels covering the reportable range
  • Reported as mean percent recovery, or as the difference from the accepted true value, with an appropriate confidence interval

Precision has three levels

Precision is not one measurement either. The three levels differ in what is allowed to vary between the repeated measurements, and they answer progressively broader questions — from whether the procedure is reproducible on one bench on one day, to whether it survives being run somewhere else entirely.

LevelWhat variesWhat it establishes
RepeatabilityNothing — same conditions, same short intervalThe procedure agrees with itself under one set of conditions
Intermediate precisionDays, analysts, equipment, environmental conditions — within one laboratoryThe procedure survives ordinary variation inside the laboratory that runs it
ReproducibilityThe laboratory itself, by inter-laboratory trialThe procedure transfers to other laboratories
Q2(R2) notes that reproducibility is usually not required for a regulatory submission, but should be considered where a procedure is being standardised — for inclusion in a pharmacopoeia, for instance, or where it will be run at multiple sites.

Specificity, and what happens when you cannot have it

Specificity is the ability to assess the analyte unequivocally in the presence of what else may be there. Where a procedure is not specific, Q2(R2) allows selectivity to be demonstrated instead: the procedure must minimise interference and show it is fit for the intended purpose. And where one procedure does not provide sufficient discrimination, the guideline recommends combining two or more. This is the formal basis for reporting chromatographic and mass data together — not belt and braces, but the documented remedy when a single measurement principle cannot separate what needs separating.

Range, linearity and the limits are one structure

These four are usually presented as separate items and are better understood as a single description of where a procedure works. The range is the interval between the lowest and highest results within which the procedure has a suitable level of response, accuracy and precision — so range is defined in terms of the other characteristics rather than alongside them. Linearity describes whether results across that range are proportional to the true values, evaluated by an appropriate statistical method such as a least-squares regression line. And the detection and quantitation limits are simply where the bottom of the range falls.

Robustness is about the real world

Robustness asks whether the procedure still works when conditions drift, and it is established deliberately rather than discovered accidentally. Q2(R2) describes robustness testing as showing the reliability of a procedure in response to deliberate variations in its parameters, along with the stability of sample preparations and reagents for the duration of the procedure where relevant. It is evaluated during development, because a procedure that only works on one instrument, on a good day, is not a procedure — it is a result that happened once.

How a validation study is put together

The characteristics are not run as an unordered checklist. They follow from the purpose, and the criteria are set before the data is generated — which is what stops a study from becoming a search for a flattering result.

  1. Intended purposeWhat the procedure will be relied upon to establish, and for what kind of sample.
  2. Characteristics selectedWhich performance characteristics are relevant to that purpose; not every procedure needs every one.
  3. Criteria set in advanceAcceptance criteria describing the numerical range, limit or desired state for each characteristic.
  4. Study conductedData generated under a protocol that states the purpose and the criteria before results exist.
  5. ConclusionWhether the evidence demonstrates the procedure is fit for that purpose — and the conclusion is bounded by it.
Validation is not permanent. Where a procedure changes, Q2(R2) contemplates partial revalidation of the performance characteristics potentially affected, rather than repeating everything or assuming nothing has moved.

What a validated method does not mean

Four inferences are commonly drawn from the word validated, and none of them follows. Understanding what the term does not carry is most of the value of understanding it at all.

  • Not that every batch was tested with it — which tests were run on a given batch is a question for the certificate
  • Not that any particular result is correct — validation establishes the procedure is capable, not that one run was error-free
  • Not that it is valid for other purposes — validation is bounded by the intended purpose it was demonstrated against
  • Not that it holds outside its range — results outside the validated interval carry none of its assurances
  • Not that it is permanent — changes to the procedure can require revalidation of the characteristics affected

Reading a report with validation in mind

The useful questions are about boundedness. A result is a measurement plus the conditions under which it can be trusted, and a report that supplies the first without the second has given you half of it.

  • Is the analytical procedure named, or only the result?
  • Is the result inside the range the procedure was established over?
  • For a low-level result, is the quantitation limit stated?
  • Are identity and quantity reported from procedures suited to each?
  • Where one procedure could not discriminate, was a second, different one used?

Research Checklist

  • Confirm batch identity and records.
  • Document all preparation inputs.
  • Keep use within controlled laboratory workflows.
  • Do not infer applied-use suitability from guide content.

Frequently Asked Questions

Does a validated method mean the result is correct?

No. Validation demonstrates that a procedure is fit for its intended purpose — that it is capable of producing reliable results under stated conditions. Whether one particular run was executed correctly is a separate question, which is why system suitability checks and the surrounding record still matter.

What is the difference between accuracy and precision?

Accuracy is closeness to the accepted true value; precision is closeness of agreement between repeated measurements of the same homogeneous sample. They are independent. A procedure carrying a systematic bias can be highly precise and consistently wrong, and its own scatter will not reveal it — which is why each is demonstrated by a different experiment.

Why would a laboratory run two different methods for one question?

Because one may not discriminate sufficiently on its own. ICH Q2(R2) allows selectivity where a procedure is not specific, and recommends combining two or more procedures where a single one does not provide sufficient discrimination. Reporting chromatographic and mass data together is that recommendation in practice.

Is validation done once?

No. It is tied to the procedure as it stands and to the purpose it was demonstrated against. Where the procedure changes, the guideline contemplates partial revalidation covering the performance characteristics the change could affect, rather than treating the original study as permanent.

Sources

The technical statements in this guide are drawn from the following. Where a definition is contested or a figure depends on method, the guide says so rather than picking one.

  1. ICH Q2(R2): Validation of Analytical ProceduresInternational Council for Harmonisation, via the European Medicines AgencyUsed for the objective of validation, the scope of procedures covered, the definition of a performance characteristic, accuracy by reference material comparison and by spiking study and its recommended data, the definition and three levels of precision, intermediate precision and reproducibility, specificity and selectivity and the recommendation to combine procedures, the definition of range, the evaluation of linearity, robustness testing, and partial revalidation.
  2. Q2(R2) Validation of Analytical Procedures: Guidance for IndustryU.S. Food and Drug AdministrationThe same guideline as adopted by the FDA.