Every point-of-care device arrives with a claim. The leaflet says the CV is under three percent, the correlation with the laboratory is 0.98, the method is certified. The service lead signs the purchase, the trainer books the sessions, the first patients are tested within a fortnight. Somewhere in that fortnight a question should have been asked and usually is not: how much error can this clinical use actually tolerate, and has anyone shown that this device, in these hands, on this site, stays inside it?

Across hospital POCT services, the diagnostics industry and software, I have seen verification treated as a formality completed after go-live, if at all, and I have seen a carefully argued specification save a service from a device that looked excellent on paper. The difference is rarely statistical sophistication. It is whether someone decided, before the first result, what "good enough" meant.

This piece is about that decision. My argument is that a point-of-care device is only verified when its measured bias and imprecision have been compared against a written specification derived from the clinical use, and that specification has to be chosen before the data are collected, not after.

8 to 23%of insulin doses were wrong in simulation with a glucose meter carrying only 5% total analytical error
0.8%desirable CV for HbA1c derived from the EFLM Biological Variation Database (half of a within-person CV of 1.6%)
1.7 to 6.2%CVs measured on four point-of-care HbA1c devices at 48 mmol/mol in one independent evaluation
25measurements over five days per level: the core of the CLSI EP15-A3 verification protocol

Where performance specifications come from

The field has twice tried to put the sources of a specification in order. The first attempt was the Stockholm consensus conference of April 1999, which set out a five-level hierarchy[1]. At the top sat the effect of analytical performance on clinical outcomes in specific settings. Below it came the effect on clinical decisions in general, judged either from biological variation or from what clinicians say they need. Then published recommendations from expert bodies, then limits set by regulators or EQA organisers, and at the bottom the current state of the art. The rule was simple: use the highest level available for the purpose.

In 2014 the EFLM met in Milan to revisit that hierarchy, and its 2015 consensus statement simplified it into three models[2]. Model 1 bases the specification on the effect of analytical performance on clinical outcomes, either directly in outcome studies or indirectly by simulating how analytical error changes clinical classifications and decisions. Model 2 bases it on the components of biological variation, aiming to keep the "analytical noise" small relative to the biological signal. Model 3 is the state of the art: the best performance technically achievable, or the performance reached by a given share of laboratories.

Stockholm 1999Milan 20141 Effect on clinicaloutcomes2 Effect on decisions:biology or clinicians3 Expert bodyrecommendations4 Regulator or EQAorganiser limits5 State of the artModel 1Clinical outcomesdirect or simulatedModel 2Biological variationwithin and between peopleModel 3State of the artbest achievable todayMilan: prefer models 1 and 2, and state the rationale, source and quality of evidence
Figure 1. The five Stockholm levels and how they map on to the three Milan models. Milan asks that preference be given to models 1 and 2 and that every specification carries a statement of its rationale, source and quality of evidence. Schematic, not measured data. Drawn from Fraser 2015 and Sandberg and colleagues 2015.

Two parts of the Milan statement matter more to a POCT service than the models themselves. First, the authors are explicit that each model has a weakness. Outcome-based specifications only work where the link between test, decision and outcome is strong. Biological variation depends on the validity of the underlying studies. State of the art is easy to find but may bear no relationship to what patients need. Second, the statement says a test with several clinical uses can carry several specifications. Its own example is glucose: one specification for critical care by simulation, another for self-monitoring in type 1 diabetes from outcome studies, a third from biological variation[2]. That is the sentence a POCT committee should have pinned to the wall. The device is not good or bad in the abstract. It is good or bad for a use.

Model 1 in practice: what outcome-based specifications look like

Randomised outcome studies of analytical quality are almost non-existent. What we have are indirect outcome studies, and the best known is a glucose simulation every POCT lead should read. Boyd and Bruns built Monte Carlo models of glucose meters with defined bias and imprecision, generated 10,000 to 20,000 pairs of true and measured glucose values for each combination, and checked whether the insulin dose chosen from the measured value matched the dose the true value called for[4].

The results were uncomfortable. A meter with a total analytical error of 5 percent produced dosing errors in roughly 8 to 23 percent of insulin doses. At 10 percent total error, 16 to 45 percent of doses were wrong. To deliver the intended dose 95 percent of the time, both bias and CV had to sit below 1 or 2 percent, depending on the glucose level and the dosing rules[4].

Set that beside the international standard for self-testing glucose systems. ISO 15197:2013 asks that at least 95 percent of results fall within 15 mg/dL of the comparison method below 100 mg/dL (about 5.6 mmol/L) and within 15 percent at or above it, and that at least 99 percent fall in zones A and B of the consensus error grid[8]. Modern meters can pass it comfortably: one recent evaluation of two systems found 97.5 to 100 percent of results within the limits and every result in zone A[8].

Both are true at once: a meter can meet the standard and, in the simulation, still steer a meaningful fraction of sliding-scale insulin doses to the wrong step. The standard was written for self-monitoring. A ward using a single reading to choose an insulin dose is a different clinical use, and Milan's own glucose example says it can deserve a different specification[2]. When a service adopts ISO 15197 as its acceptance criterion for a hospital meter without asking which use it is specifying for, it has quietly chosen a model and a risk without writing either down.

The device is not good or bad in the abstract. It is good or bad for a use, and the specification is where that use is written down.

Model 2: what biology says, and why point-of-care devices struggle with it

The biological variation model rests on simple physiology. A person's HbA1c, creatinine or INR fluctuates around a homeostatic set point, and that within-subject variation is expressed as CVI. Different people sit at different set points, and the spread between them is the between-subject variation, CVG. If the analyser's own noise is small relative to CVI, it adds little to the variation a clinician is already living with. If its bias is small relative to the combined spread of CVI and CVG, a population reference interval or decision limit still classifies people correctly.

From that reasoning come the formulas the EFLM Biological Variation Database applies to every measurand it holds[3]. Desirable imprecision is a CV below half the CVI. Desirable bias is below a quarter of the square root of CVI squared plus CVG squared. A desirable total allowable error combines the two as bias plus 1.65 times the CV. Minimum specifications loosen the multipliers to 0.75 and 0.375, optimum specifications tighten them to 0.25 and 0.125. Importantly, it publishes meta-analysed estimates with confidence intervals rather than a figure from one study.

Run those formulas for four analytes commonly tested at the point of care and the problem is plain.

AnalyteCVICVGDesirable CVDesirable biasDesirable total errorPublished POC CV
HbA1c (IFCC units)1.6%7.3%0.8%1.9%3.2%1.7 to 6.2% (4 devices)[6]
Glucose4.7%8.0%2.3%2.3%6.2%up to 2.6% intermediate, 4.2% repeatability[8]
INR (healthy subjects)2.5%4.6%1.3%1.3%3.4%3.6% (1 system)[9]
Creatinine4.4%15.8%2.2%4.1%7.7%5.8 to 11.3% (1 device)[10]

The CVI and CVG values are the database's meta-analysed estimates as retrieved on 1 October 2026; the specifications are calculated from them with the database's own formulas[3]. The device CVs come from four separate published evaluations and are not head-to-head. The population matters too. The database INR estimates come from healthy subjects; in 245 stable patients on long-term warfarin, mean within-person variation was 9.0 percent, giving a desirable imprecision of below 4.5 percent for the 2.0 to 3.0 therapeutic range[24]. Against that monitoring specification the published point-of-care INR CV of 3.6 percent passes.

0%2%4%6%8%10%12%Imprecision (CV, %)HbA1c4 POC devicesGlucosemeters: IP, repeatabilityINRhealthy-subject limitswarfarin limit 4.5%Creatinine1 POC device, rangeDesirable (0.5 x CVI)Minimum (0.75 x CVI)Published POC CV
Figure 2. Imprecision specifications from biological variation against published point-of-care CVs. Every plotted HbA1c, INR and creatinine device exceeds the minimum limit derived from healthy-subject biological variation; for INR, the warfarin-derived limit of 4.5 percent is also marked. These limits alone do not establish clinical unfitness. Source: EFLM Biological Variation Database (CVI estimates, accessed 1 October 2026); device CVs from Lenters-Westra and English 2018, Pleus and colleagues 2024, Tait and colleagues 2019, Nataatmadja and colleagues 2020; warfarin limit from van den Besselaar and colleagues 2012.

For HbA1c the gap is stark. The within-person variation of HbA1c is tiny, around 1.6 percent in the current database estimate, with a confidence interval of 1.3 to 2.5 percent, because it integrates glucose exposure over the lifespan of a red cell[3]. Half of that is 0.8 percent. The best of four point-of-care HbA1c devices in an independent evaluation following CLSI EP-5 and EP-9 protocols achieved 1.7 percent at 48 mmol/mol, and the worst 6.2 percent[6]. The laboratory fares little better. When the IFCC task force on HbA1c standardisation tested the biological variation model on external quality assessment data, 48 percent of individual laboratories and none of 26 instrument groups met even the minimum criterion[5].

That finding is what pushed the IFCC to a different answer. Its task force recommended a sigma-metrics model with a default total allowable error of 5 mmol/mol at an HbA1c of 50 mmol/mol, which is 10 percent, chosen because it reflects the difference between two consecutive results that would prompt a change of therapy and the gap between the upper end of low risk and the diagnostic limit[5]. Under that model, 77 percent of laboratories and 12 of 26 instrument groups reached the 2 sigma criterion recommended for routine use. The same 10 percent target is the yardstick in the IFCC certification that manufacturers advertise, and an independent evaluation published in 2025 found that only 5 of 19 point-of-care HbA1c devices met both IFCC and NGSP criteria[7].

The 2015 task force paper used a CVI for HbA1c of 2.77 percent in IFCC units[5]. The current database estimate, meta-analysed from two studies, is 1.6 percent[3]. Biological variation estimates move as the evidence improves, which means a specification derived from them should carry its source and date.

Total error, measurement uncertainty and the argument about sigma

How to compare a device against a specification has been argued over for a decade. The total error tradition, associated with Westgard and colleagues, adds bias and imprecision into one error budget and asks whether a result is likely to fall within a total allowable error. The sigma metric falls out of it naturally: the allowable error minus the absolute bias, divided by the CV, all in percent[13]. On that scale six sigma is world class, five excellent, four good, three marginal, two poor, and below two unacceptable, and the sigma value drives how many controls a method needs and how often[13].

The metrological tradition approaches the same problem differently. Measurement uncertainty, as codified for medical laboratories in ISO/TS 20914:2019, expresses the dispersion of values that could reasonably be attributed to the measurand, assumes known bias is corrected or included as an uncertainty component, and is built largely from long-term internal QC data and calibrator uncertainty. Its scope explicitly includes quantitative values produced near medical decision thresholds by point-of-care testing systems[17]. ISO 15189:2022, which now absorbs the point-of-care requirements previously held in the withdrawn ISO 22870[18], requires that measurement uncertainty be evaluated and maintained for its intended use, where relevant, and compared against performance specifications[19].

Oosterhuis and Theodorsson put the tension plainly: the total error theory dominates clinical chemistry practice but is not accepted in other fields of metrology, while uncertainty theory suffers from complex mathematics and a perceived impracticability, and the way forward should be evolution rather than revolution[16]. The sharper critique is of sigma itself. Oosterhuis and Coskun showed that if a method sits exactly at the desirable biological variation specification, the conventional sigma calculation simply returns the coverage factor of 1.65, whatever the method is really doing[14]. They proposed a sigma based on the ratio of CVI to analytical CV instead. Westgard and colleagues replied that the alternative ignores bias, which for point-of-care devices calibrated outside the laboratory's control is often the dominant error[15].

Worked example

Take a hypothetical point-of-care HbA1c device verified at 50 mmol/mol with a bias of +1.0 mmol/mol (2.0 percent) and a CV of 2.4 percent. Against the IFCC total allowable error of 10 percent[5], sigma is (10 minus 2.0) divided by 2.4, which is 3.3: comfortably above the IFCC's 2 sigma line for routine use, marginal on the Westgard scale. Now judge the same data against the biological variation target from the table above. Desirable bias is 0.25 times the square root of (1.63 squared plus 7.3 squared), which is 1.87 percent, desirable CV is 0.82 percent, and desirable total error is 1.87 plus 1.65 times 0.82, about 3.2 percent. Sigma becomes (3.2 minus 2.0) divided by 2.4, which is 0.5. Nothing about the device changed. And a device sitting exactly on the desirable bias and CV would score (3.21 minus 1.87) divided by 0.815, which is 1.65, the coverage factor, as Oosterhuis and Coskun predicted.

0123456Sigma metric = (allowable error minus |bias|) / CVunacceptablepoormarginalgoodexcellentFour POC HbA1c devices, allowable error 10% (published)D1.4C2.1B4.0A5.8biology target 3.2%: 0.5IFCC target 10%: 3.3Bottom markers: one hypothetical device (bias 2.0%, CV 2.4%), two choices of target
Figure 3. Four published point-of-care HbA1c devices on the sigma scale, judged against a 10 percent allowable error, and the hypothetical device from the worked example scored against two different targets. Source: Lenters-Westra and English, Journal of Diabetes Science and Technology, 2018 (sigma 5.8, 4.0, 2.1, 1.4); worked example calculated from EFLM and IFCC targets.

For a small service my view is pragmatic: the debate does not change what you do on the bench. Write the specification model down first. Verify bias and imprecision separately against their own limits, because one total error figure can hide a large bias behind a small CV. Use sigma, alongside the error detection of your QC rules, workload, clinical risk and the manufacturer's minimum requirements, to plan QC; it is only as meaningful as the allowable error you fed it. And estimate measurement uncertainty from long-term QC, because ISO 15189:2022 expects it.

Serial results: why the reference change value matters more at the point of care

Much point-of-care testing is monitoring: HbA1c in the diabetes clinic, INR in the anticoagulation clinic, creatinine before and after a new drug. The question is whether a result has changed. That question has a formal answer, the reference change value. In Fraser's formulation, RCV equals the square root of 2, times Z, times the square root of CVA squared plus CVI squared, where Z reflects the probability required[12]. With Z at 1.96 for a two-sided 95 percent probability, the multiplier is about 2.77.

The formula makes the point-of-care trade-off visible. CVI is fixed by physiology. CVA is the only term the service controls, and point-of-care devices often have larger CVs than the laboratory analysers they stand in for. The classic formula is symmetric; the EFLM database now also presents asymmetric RCVs for increases and decreases[3]. The symmetric version is used below because it can be checked with a calculator.

Worked example

A patient's HbA1c falls from 53 to 49 mmol/mol on a point-of-care device with a CV of 2.4 percent. With the EFLM CVI of 1.63 percent, RCV is 2.77 times the square root of (1.63 squared plus 2.4 squared), which is 2.77 times 2.90, or 8.0 percent: about 4.3 mmol/mol at 53. The observed fall of 4 mmol/mol (7.5 percent) does not exceed this estimated RCV; a genuine change remains possible, but the result alone cannot show one. Now creatinine. With a CVI of 4.39 percent and the CVs one published evaluation measured at 73 and 117 micromol/L, 5.8 to 11.3 percent[10], RCV is 20.2 to 33.6 percent. Assuming those CVs hold at a baseline of 100 micromol/L, that is a rise of about 20 to 34 micromol/L before a change is statistically clear. NICE detects acute kidney injury on a rise of 26 micromol/L or more within 48 hours[11]. The RCV does not validate or invalidate that criterion, but it shows how much a noisy device can blur it.

0%5%10%15%20%25%30%35%Change needed between two results to be 95% confident it is realHbA1c4.5% biology only6.5% with CVA 1.7%10.5% with CVA 3.4%INR (warfarin)24.9% biology only26.9% with CVA 3.6%Creatinine12.2% biology only20.2% with CVA 5.8%33.6% with CVA 11.3%Within-person variation onlyPlus published POC imprecision
Figure 4. Reference change values at 95 percent probability, from within-person biological variation alone and with published point-of-care imprecision added. For creatinine, adding device imprecision raises the RCV from 12.2 percent to 20.2 to 33.6 percent, about 1.7 to 2.8 times the biological value. For INR in patients on warfarin, biology dominates and the device adds little. Source: CVI from the EFLM Biological Variation Database (HbA1c, creatinine) and van den Besselaar and colleagues 2012 (INR on warfarin); CVA from Lenters-Westra and English 2018, Tait and colleagues 2019, Nataatmadja and colleagues 2020; RCV by the formula in Fraser 2012.

None of this says point-of-care creatinine or HbA1c should not be used for monitoring. It says a monitoring specification has to be set with the RCV in mind, and clinicians need to know how large a change the device can detect. The same evaluation that reported those creatinine CVs found a mean positive bias of 12.7 micromol/L against an IDMS-traceable laboratory method and misclassification across all CKD stages[10], which is a reminder that bias and imprecision act together on any decision with a fixed threshold.

How verification is actually designed

The manufacturer has validated the device; the service has to show that the claims hold on its site, with its operators and samples, and that they satisfy its own specification. Two CLSI documents carry most of the weight.

Precision and bias: EP15-A3

CLSI EP15-A3 is built for exactly this situation and is explicitly intended to be usable in point-of-care and physician office settings as well as large laboratories[20]. Its core design is five replicates of each sample per run over at least five days, giving 25 or more results for each of two or more materials at different medical decision concentrations, from which repeatability and within-laboratory precision are estimated and compared with the manufacturer's claim. The same experiment, run on materials with known or assigned values, gives an estimate of bias. Choose those levels near the decision limits that matter for your clinical use. The statistical point that trips services up is that a verification does not fail simply because your observed CV is a little above the claim; EP15 compares the observed value with a verification limit calculated from the claimed value and the size of your experiment. Equally, a CV that passes the manufacturer's claim has not yet been compared with your specification. Those are two separate tests.

Method comparison: EP09, Bland-Altman and Passing-Bablok

Bias against the laboratory, on real patient samples, is the domain of CLSI EP09. It covers difference plots, Deming and Passing-Bablok regression, and bias at decision points with confidence intervals[21]. Published POCT evaluations give a sense of scale: the HbA1c study above used 40 fresh patient samples on two instruments, eight a day over five days, measured in duplicate[6]. Samples should span the measuring interval, weighted towards decision limits.

Two analyses answer different questions. The Bland-Altman difference plot shows the mean difference and the limits within which 95 percent of differences between the two methods lie, and its authors were blunt that correlation coefficients, still common in vendor brochures, are misleading for agreement[22]. A correlation of 0.98 can coexist with a clinically important bias. Passing-Bablok regression estimates slope and intercept without assuming a particular distribution of the data or errors, and its confidence intervals tell you whether proportional bias (slope differing from 1) or constant bias (intercept differing from 0) is more than chance[23]. Use both. The difference plot shows whether agreement is clinically acceptable; the regression shows what kind of bias you have and whether it depends on concentration.

What sinks a verification in practice

In verification work I have seen at small private clinics, the studies that failed rarely failed on the analytical numbers. They failed before the analyser got a fair hearing. A high repeat rate on urinalysis negatives at first dip, which said more about technique than the strips. Test strips not fully immersed in the control solution, so the control run measured the operator rather than the device. The wrong QC material for the device in use. A whole-blood sample labelled as serum, which turned a method comparison into a matrix comparison. And an analyser that could not distinguish a QC run from a patient run, so that QC results had to be transcribed by hand into a spreadsheet, with every transcription a fresh opportunity for error and no reliable audit trail. Poor immersion or handling can inflate a CV, but identification errors, wrong materials and transcription failures can pass a precision study untouched. Every one of them would reach a patient result.

A verification plan a small service could actually run

The following is a framework for what governance should ask for, not a substitute for the manufacturer's instructions or a locally validated SOP. Where they differ, those take precedence. It also assumes the wider POCT governance MHRA describes is in place: laboratory involvement, recorded operator competence, IQC, EQA participation and incident reporting through the Yellow Card scheme[25]. Analytical verification does not establish regulatory conformity or accreditation.

  1. Write the intended use in one sentence. Diagnosis, monitoring, screening, or triage; which patients; which decision the result will drive and at what concentration. Everything else follows from this.
  2. Choose the specification model and record why. Use an outcome-based or consensus specification where one exists for that use (the IFCC 10 percent total allowable error for HbA1c, ISO 15197:2013 for self-monitoring glucose, with its limits noted), otherwise the EFLM biological variation specifications at desirable or, with a stated reason, minimum level. Record the source, the version or access date, and separate limits for bias and imprecision.
  3. Check the claims against the specification before purchase. If the manufacturer's own precision and bias claims cannot meet your specification, verification on site will not rescue it. This is the cheapest point to say no.
  4. Run a precision and trueness experiment. Follow an EP15-style design: two levels near the decision limits, five replicates a run over at least five days, run by the routine operators. For trueness, use material with a traceable assigned value that behaves like patient samples, or rely on the patient-sample comparison below; a value assigned by the same method can hide a shared calibration bias.
  5. Run a method comparison on patient samples. Aim for around 40 paired samples spread across the measuring interval and concentrated near decision limits, with the correct sample type at both ends and the laboratory sample handled to the laboratory's own requirements. Analyse with a difference plot and Passing-Bablok regression; report bias with confidence intervals at each decision limit.
  6. Audit the pre-analytical steps during the study. Observe technique, sample labelling, control handling and how results are recorded. Log every repeat and its reason. A high repeat rate is a finding, not noise.
  7. Compare, decide and sign. Set observed bias and imprecision against the written limits, compute the sigma metric against the chosen allowable error, and record a decision: accept, accept with conditions, or reject. Name who signed it.
  8. Carry the numbers into service. Use the verified CV to calculate the reference change value for monitoring uses and tell the clinicians what change the device can detect. Use the sigma value, with QC rule performance, workload, risk and the manufacturer's requirements, to plan QC frequency. Start building a measurement uncertainty estimate from long-term QC, and confirm that QC and patient results can be told apart in the record without manual transcription.

A service doing this for one analyte will find most of the effort is in the first two steps, as it should be. Once the use and the specification are written down, the statistics are a recipe, and the conversation with clinicians changes. Instead of "the device is accurate", the service can say what error it carries, why that error is acceptable for this decision, and how large a change between two results means something. That sentence, written before the first patient is tested, is what "good enough" looks like.

Sources and notes

Biological variation estimates are the EFLM database's meta-analysed values as retrieved on 1 October 2026 and will change as studies are added; the specifications in the table and Figure 2 are calculated from them with the database's own formulas. Published device CVs come from four separate evaluations with different designs, concentrations and sample types and are not a head-to-head comparison; devices are not named in the figures because the purpose is to show the size of the gap, not to rank products. Sigma values in Figure 3 are those published by the study authors against a 10 percent allowable error; the worked example device is hypothetical. Reference change values in Figure 4 and the worked examples use the symmetric two-sided formula at 95 percent; asymmetric log-normal RCVs will differ slightly, and the creatinine example assumes CVs measured at 73 and 117 micromol/L hold at 100. INR biological variation differs sharply between healthy subjects and patients on warfarin. The verification observations are drawn from the author's own work and are deliberately unattributed and unquantified. Figure 1 is a schematic; Figures 2 to 4 plot sourced numbers. CLSI and ISO documents are cited from their publishers' summaries.

  1. Fraser CG. The 1999 Stockholm Consensus Conference on quality specifications in laboratory medicine. Clinical Chemistry and Laboratory Medicine, 2015 (the five-level hierarchy).
  2. Sandberg S, Fraser CG, Horvath AR, et al. Defining analytical performance specifications: Consensus Statement from the 1st Strategic Conference of the European Federation of Clinical Chemistry and Laboratory Medicine. Clinical Chemistry and Laboratory Medicine, 2015 (three models; preference for models 1 and 2; glucose example of multiple specifications).
  3. European Federation of Clinical Chemistry and Laboratory Medicine. EFLM Biological Variation Database. Accessed 1 October 2026 (CVI and CVG: HbA1c IFCC 1.63 and 7.3%; glucose 4.67 and 8.04%; creatinine 4.39 and 15.81%; INR 2.5 and 4.59%; specification formulas).
  4. Boyd JC, Bruns DE. Quality specifications for glucose meters: assessment by simulation modeling of errors in insulin dose. Clinical Chemistry, 2001 (8 to 23% dose errors at 5% total error; 16 to 45% at 10%).
  5. Weykamp C, John G, Gillery P, et al. Investigation of 2 models to set and evaluate quality targets for Hb A1c: biological variation and sigma-metrics. Clinical Chemistry, 2015 (TAE 5 mmol/mol at 50 mmol/mol; 48% of laboratories and 0 of 26 groups met BV minimum; 77% and 12 of 26 met 2 sigma).
  6. Lenters-Westra E, English E. Evaluation of four HbA1c point-of-care devices using international quality targets: are they fit for the purpose? Journal of Diabetes Science and Technology, 2018 (CVs 1.7, 2.4, 3.4, 6.2% at 48 mmol/mol; sigma 5.8, 4.0, 2.1, 1.4; 40 samples over 5 days).
  7. Lenters-Westra E, Singh P, Vetter B, English E. Challenges in HbA1c point-of-care testing: only 5 of 19 HbA1c point-of-care devices meet IFCC and NGSP certification criteria on independent evaluation. Clinical Chemistry, 2025.
  8. Pleus S, Jendrike N, Baumstark A, et al. Evaluation of system accuracy, precision, hematocrit influence, and user performance of two blood glucose monitoring systems based on ISO 15197:2013/EN ISO 15197:2015. Diabetes Therapy, 2024 (ISO 15197 criteria; repeatability CV up to 4.2%, intermediate precision up to 2.6%).
  9. Tait RC, Hung A, Gardner RS. Performance of the LumiraDx Platform INR test in an anticoagulation clinic point-of-care setting compared with an established laboratory reference method. Clinical and Applied Thrombosis/Hemostasis, 2019 (mean CV 3.60% across three strip lots).
  10. Nataatmadja M, Fung AWS, Jacobson B, et al. Performance of StatSensor point-of-care device for measuring creatinine in patients with chronic kidney disease and postkidney transplantation. Canadian Journal of Kidney Health and Disease, 2020 (CV 5.8 to 11.3%; mean bias 12.7 micromol/L).
  11. National Institute for Health and Care Excellence. Acute kidney injury: prevention, detection and management (NG148). NICE, 2019 (rise of 26 micromol/L or more within 48 hours, recommendation 1.3.1).
  12. Fraser CG. Reference change values. Clinical Chemistry and Laboratory Medicine, 2012 (RCV formula).
  13. Westgard S, Bayat H, Westgard JO. Analytical Sigma metrics: a review of Six Sigma implementation tools for medical laboratories. Biochemia Medica, 2018 (sigma equation and scale).
  14. Oosterhuis WP, Coskun A. Sigma metrics in laboratory medicine revisited: we are on the right road with the wrong map. Biochemia Medica, 2018 (sigma equals the 1.65 coverage factor at desirable specifications).
  15. Westgard S, Bayat H, Westgard JO. Mistaken assumptions drive new Six Sigma model off the road. Biochemia Medica, 2019.
  16. Oosterhuis WP, Theodorsson E. Total error vs. measurement uncertainty: revolution or evolution? Clinical Chemistry and Laboratory Medicine, 2016.
  17. International Organization for Standardization. ISO/TS 20914:2019, Medical laboratories, practical guidance for the estimation of measurement uncertainty (scope includes POCT values near decision thresholds).
  18. Eurachem. ISO 15189:2022, a new task for medical laboratories. Eurachem leaflet (POCT requirements incorporated; ISO 22870 withdrawn; measurement uncertainty and ISO/TS 20914).
  19. International Organization for Standardization. ISO 15189:2022, Medical laboratories, requirements for quality and competence (clause 7.3.4, evaluation of measurement uncertainty).
  20. Clinical and Laboratory Standards Institute. EP15-A3, User verification of precision and estimation of bias. CLSI, 2014 (five-day protocol; 25 or more measurements per material; suitable for point-of-care settings; verification limit).
  21. Clinical and Laboratory Standards Institute. EP09, Measurement procedure comparison and bias estimation using patient samples. CLSI, third edition.
  22. Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet, 1986.
  23. Passing H, Bablok W. A new biometrical procedure for testing the equality of measurements from two different analytical methods. Journal of Clinical Chemistry and Clinical Biochemistry, 1983.
  24. van den Besselaar AM, Fogar P, Pengo V, et al. Biological variation of INR in stable patients on long-term anticoagulation with warfarin. Thrombosis Research, 2012 (245 patients; mean within-person CV 9.0%; desirable imprecision below 4.5%).
  25. Medicines and Healthcare products Regulatory Agency. Management and use of IVD point of care test devices. GOV.UK, updated 2026 (laboratory involvement, competence, IQC, EQA, Yellow Card reporting).