B9 · Evidence-Based Practice & Critical Appraisal
> Currency and provenance — 62 references · median 2010, range 1986-2024, 3 % from 2022 on · provenance: verified external 89 % (55) · MEDLIB corpus 11 % (7).
Domain B · Patient Assessment & Consultation · The governing chapter: the [A–D] tag on every figure, dose and interval in this atlas is assigned by the rules defined here. If a chapter contradicts B9, the error is in the chapter.
> Tags: [A] guideline/consensus/SR with explicit year · [B] primary literature with verified PMID or DOI · [C] reference monograph or textbook · [D] slide, course or opinion, never sufficient alone · [MEDLIB] own document corpus · [MODELO] structure and wording only, never a figure · (P) model reasoning, never a dose · ⚠ disputed or drifting figure.
Subchapters
- [x] B9.1 · In 30 seconds
- [x] B9.2 · Applicable normative framework (ES and EU)
- [x] B9.3 · The step-by-step procedure
- [x] B9.4 · Templates and documents
- [x] B9.5 · Frequent errors and their cost
- [x] B9.6 · Metrics: what is measured and its reference value
- [x] B9.7 · Spanish particularity
- [x] B9.8 · Organizational alternatives
B9.1 · In 30 seconds
The thesis, unsoftened: aesthetic medicine is a discipline with very little high-quality evidence and an enormous volume of literature. Most of what is published is case series, open-label studies, expert consensus and trials funded by the maker of the product under test. That does not invalidate the field; it imposes a different discipline from the rest of medicine: know, at every moment, how firmly you are holding what you are doing, and say so [1][9]. The operational consequence: in this atlas no claim ships without a tag. If a sentence carries no grade, it is not finished.
The evidence ladder, top to bottom [A]: systematic review/meta-analysis of RCTs > individual RCT > cohort > case-control > case series > mechanism-based reasoning / uncontrolled expert opinion [2][4]. Most strong aesthetic recommendations live at the bottom two rungs with a strong recommendation carried by consensus and plausibility. The canonical example is high-dose pulsed hyaluronidase for filler vascular occlusion: nobody will ever run the RCT, and the recommendation is still as strong as a recommendation can be. GRADE names why that is coherent: it separates certainty of evidence (high / moderate / low / very low) from strength of recommendation (strong / conditional) [3][5].
The numbers and red lines to hold at the chairside:
| Quantity | Reference value / red line | Where it bites | Ref |
|---|---|---|---|
| p value | 0.05 is an arbitrary convention, not a truth threshold; p<0.05 says the data are unlikely under the null, nothing about effect size |
Large n makes trivial differences "significant" | [35] |
| 95% CI | crosses 1 (ratios) or 0 (differences) → not significant; wide CI = the study knows little | Read the CI, not the point estimate | [36] |
| Effect size / MCID | ask the Minimal Clinically Important Difference of the scale; no MCID published → the scale cannot prove usefulness | 0.3 points on a 5-point scale can be robust and invisible | [39] |
| ARR vs RRR | demand the absolute risk; "reduces risk 50 %" can mean 0.02 % → 0.01 % | The commonest marketing inflation | [37][38] |
| NNT / NNH | lower NNT is better; always pair with baseline risk and time horizon; NNH is directly usable for complications | Translates "0.3%" into "1 in 333" | [37] |
| Heterogeneity I² | 0–40% may be unimportant; 30–60% moderate; 50–90% substantial; 75–100% considerable | I²>50% → pooling questionable | [40] |
| Dropout | losses >20 % threaten validity; per-protocol overstates effect vs intention-to-treat | The dissatisfied leave first | [44] |
| Follow-up | read it before the safety line; late nodules/granuloma appear at 6–24 months | A 12-week study cannot see them | [10] |
| Sample size | n=10–20 detects only huge effects; "no significant difference" ≠ "no difference" | Underpowered null is uninformative | [9] |
| Diagnostic LR | LR+ >10 or LR− <0.1 = large shift; 5–10 / 0.1–0.2 = moderate; PPV/NPV move with prevalence | ROC, sens/spec, likelihood ratios | [42][43] |
| Sponsorship | industry-funded studies: favorable efficacy RR 1.27 (1.17–1.37), favorable conclusions RR 1.34 (1.19–1.51) | An industry bias standard risk-of-bias tools do not catch | [47] |
The five-step reflex (the 5A cycle) [A]: Ask (frame it as PICO) → Acquire (technical sheet/IFU first, then guideline, then SR, then primary) → Appraise (design, risk of bias, size, follow-up, conflicts) → Apply (to this patient, phototype, severity, injector) → Assess (audit your own outcomes) [1]. The most profitable single filter in this field: look at the follow-up before the result.
Stable vs changing claims, the distinction that avoids most errors: anatomy, biochemical mechanism and basic pharmacology are stable (a 2009 source still serves, if the year is visible). Regulation, product availability, the technical sheet, the recommended dose, device parameters and complication management are changing (a 2009 source does not serve, however good). Grade currency by which category the claim is in.
| Claim class | Currency bar | Carried-forward model estimate |
|---|---|---|
| Stable (anatomy, mechanism, basic pharmacology) | a dated old source still serves | applying this stable-vs-changing split first heads off, by the prior edition's estimate, on the order of ~80 % of the field's avoidable appraisal errors (a pedagogical figure, not a sourced datum) |
| Changing (regulation, IFU/dose, device parameters, complication protocols) | needs a current source | the source that ages fastest; recheck against CIMA/IFU quarterly |
Classic trap: reading disposition: answer or a slide title as evidence. A course deck ([D]) tells you where to look; it never closes a clinical question, and a dose resting on a single UPO slide is never_sufficient_alone.
The ladders, in full
Oxford OCEBM 2011 Levels of Evidence [A] · a fast, point-of-care heuristic that grades not only therapy but prevalence, diagnosis, prognosis and harm [2]:
| Level | Therapy question answered by | Aesthetic reality |
|---|---|---|
| 1 | SR of RCTs, or n-of-1 trial | Almost none exists for injectables/EBD |
| 2 | Individual RCT, or observational study with dramatic effect | Split-face RCTs of toxin/filler; small, short |
| 3 | Non-randomised controlled cohort / follow-up study | Most "comparative" filler papers |
| 4 | Case series, case-control, historically controlled | The bulk of the field |
| 5 | Mechanism-based reasoning | "It makes biological sense" |
OCEBM grades recommendations A–C and adds a rule OCEBM authors state explicitly: a level can be downgraded for study quality, imprecision, indirectness, or because the effect size is small; and upgraded for a large or dramatic effect [2]. Murad's revision of the pyramid is the current mental model: the lines between layers are wavy, not straight, because study quality can move a study up or down, and the systematic review is not a layer at all but the lens through which the layers are read [4].
GRADE [A] is the system serious guidelines use. Its contribution is to hold apart two things the classical hierarchy blurs [3][5][6]:
| Axis | Levels | Meaning |
|---|---|---|
| Certainty of evidence | High · Moderate · Low · Very low | How confident we are the true effect lies near the estimate |
| Strength of recommendation | Strong · Conditional (weak) | How confidently we tell a clinician to act |
Certainty starts high for RCT bodies, low for observational bodies, then moves:
- Downgrade (5 domains): risk of bias · inconsistency (unexplained heterogeneity) · indirectness (population, intervention, comparator or outcome not the one you face) · imprecision (wide CI, few events) · publication bias [6].
- Upgrade (3 domains, for observational data): large magnitude of effect · dose–response gradient · all plausible residual confounding would reduce the observed effect [6].
The move from certainty to a recommendation weighs four more things: balance of benefits vs harms, certainty itself, patient values and preferences, and resource use; the Evidence-to-Decision framework makes each explicit and auditable [7][8]. This is why a strong recommendation on low certainty is legitimate: hyaluronidase in occlusion is low-certainty, strong-recommendation, and correctly so.
The atlas [A–D] tags and their correspondence:
| Tag | Meaning here | Approx. level | Rule |
|---|---|---|---|
[A] |
Guideline, consensus or SR with explicit year | OCEBM 1–2, or formal consensus | Currency stamped |
[B] |
Primary literature with verified PMID/DOI | 1–4 by design | Identifier must resolve, not be recalled |
[C] |
Reference monograph / textbook chapter | 4–5 | Stable-claim lane |
[D] |
Slide, course, opinion. Never alone | 5 | Points where to look; closes nothing |
[MEDLIB] |
Own document corpus | Variable; open the file | Provenance, not automatic authority |
[MODELO] |
Structure, wording, organisation | Not evidence | Never a figure |
(P) |
Model reasoning | Not evidence | Never a dose, ratio or published identifier |
| ⚠ | Disputed or drifting figure | · | Two values kept, never averaged |
Four usage rules that make the system mean something: (1) [B] requires a real, verified identifier: a recalled one is not [B], and mis-recalled identifiers are a documented failure mode in this project, not a hypothesis. (2) [MODELO] and (P) never carry a dose, concentration, depth or interval: if the model is the only thing holding a figure, the figure is not printed. (3) [D] never travels alone. (4) Conflicts are printed whole, with their scenario, never averaged: if one source says 1:2.5 and another 1:3, both are written, with why they differ.
GRADE vs OCEBM: the grounded divergence [A]. The two systems classify the same study differently, and the atlas must not blend them [3][2]:
Consensus: both put SR-of-RCTs at the top and mechanism-based reasoning at the bottom, and both formally downgrade for risk of bias, imprecision and indirectness [2][6].
Discrepancy: an OCEBM Level 2 / Grade B RCT can be GRADE low certainty once inconsistency and imprecision are applied; OCEBM answers "what design is this, fast?" while GRADE answers "how much should a guideline trust this body of evidence?" · decide by use: point-of-care triage of a single paper → OCEBM level as shorthand; a claim that will become a chapter recommendation → GRADE certainty. This atlas assigns [A]/[B] by GRADE certainty for anything that drives a recommendation and keeps the OCEBM level only as a design shorthand; the two labels are never merged into one number [3][6].
Worked example: grading the atlas's own strongest low-evidence claim [A]. Take "give high-dose pulsed hyaluronidase for filler-induced vascular occlusion". Walk GRADE [3][6][7]:
- Body of evidence: case reports and small series, no RCT and none possible (you cannot randomise an occluding eye to placebo). Certainty starts low (observational) and there is serious imprecision → very low.
- But three upgrade signals fire: the effect is large and dramatic (vision or skin saved vs lost), there is a dose–response logic (more enzyme, faster clot lysis), and all plausible confounding would work against the observed benefit (the worst cases get the most enzyme). Certainty is still low, but the direction is not in doubt.
- Evidence-to-Decision: benefit (tissue/vision) vastly outweighs harm (transient swelling, rare allergy), patient values are unanimous, the resource is cheap. → Strong recommendation, low certainty.
This is the shape of most defensible aesthetic recommendations: the certainty tag is honest ([A] consensus, low certainty) and the action is still firm. A reader who demanded an RCT before acting would be paralysed, which is the fundamentalist error of B9.8.
How the tag is assigned, sampled across chapters [MODELO]:
| Claim type | Typical tag | Currency bar |
|---|---|---|
| "The facial artery runs here" (anatomy) | [C] monograph |
Stable; a 2009 atlas still serves |
| "Toxin onset at X days" (pharmacology) | [A]/[B] with year |
Semi-stable; recheck at revision |
| "Filler brand Z lasts N months" (product) | [B] + IFU [A] |
Changing; recheck the IFU quarterly |
| "Manage occlusion with hyaluronidase" (protocol) | [A] consensus |
Changing; follow the latest consensus |
| "This looks better" (aesthetic judgement) | (P) or [D] |
Never carries a number |
The voices, and the one that carries a safety rule [MODELO]. The atlas dropped the (S) corpus-source and (E) external-source markers because [C]/[D] and [A]/[B] already encode them. Only (P) survives, because nothing else marks a sentence as the model's own reasoning, and that distinction carries a hard rule: a (P) sentence never states a dose, a ratio or a published identifier. If the reasoning needs a number, the number must come from a tagged source, not from the model.
The one-line summary of the whole chapter (P): in aesthetic medicine you will usually act on low evidence, and that is legitimate; the discipline is to know the grade, keep the conflicting values whole, read the follow-up before the safety line, demand the absolute number, and write down why you did what you did. Everything below is the detail of those five moves.
Classic trap: citing "it is proven" from a Level 4–5 base. The professional distinction is not acting on low evidence (often obligatory) but acting without knowing the evidence is low, and not telling the patient so [1][9].
B9.2 · Applicable normative framework (ES and EU), with the cited norm
Two bodies of "norm" govern evidence in aesthetic medicine, and they are different things: reporting/methodology standards (how a study must be written and appraised, from EQUATOR and GRADE) and legal/regulatory law (how evidence may be generated and what may be claimed, from EU and Spanish statute). A clinician who confuses "the trial did not follow CONSORT" (a reporting defect) with "the product is not authorised" (a legal defect) misreads both.
The norm map, by domain [A]:
| Domain | Governing norm | What it demands | Ref |
|---|---|---|---|
| RCT reporting | CONSORT 2010 (25 items + flow diagram) | Pre-specified primary outcome, randomisation, blinding, ITT, harms | [14] |
| Trial protocol | SPIRIT 2013 | A registered, dated protocol before enrolment | [19] |
| Harms reporting | CONSORT-Harms 2004 | Adverse events as a named, quantified outcome | [20] |
| SR / meta-analysis | PRISMA 2020 (27 items + flow diagram) | Search reproducible, selection auditable, bias assessed | [15][16] |
| Observational | STROBE (22 items) | Confounding, selection, missing data made explicit | [17] |
| Diagnostic accuracy | STARD 2015 (30 items) | Reference standard, spectrum, indeterminate results | [18] |
| Guideline appraisal | AGREE II (23 items, 6 domains) | Rigour of development, editorial independence | [25] |
| Research ethics | Declaration of Helsinki 2013 · ICH-GCP | Consent, ethics committee, favourable risk/benefit | [26] |
| Trial registration | ICMJE requirement | Prospective public registration or no publication | [27] |
| Medicines (toxin) | EU CTR 536/2014 · Spanish RD 1090/2015 · AEMPS | CEIm approval, CTIS submission, authorised indication | [29][30] |
| Devices (filler, threads, most EBD) | EU MDR 2017/745 | Clinical evaluation, CE mark, post-market surveillance | [28] |
| Off-label / special use | Spanish RD 1015/2009 | Documented justification, informed consent, no routine use | [31] |
| Biomedical research | Spanish Ley 14/2007 | Legal frame for human-subject and sample research | [32] |
| Guideline repository | EQUATOR Network · GuiaSalud (SNS) | Where to find the applicable reporting guideline | [33][55] |
EQUATOR reporting guidelines are the "cited norm" for reading a paper [A]. Each has a matching flow or checklist; the two that a clinician meets most are the PRISMA and CONSORT flow diagrams. A systematic review that cannot show a PRISMA flow (how many records found, screened, excluded and why, included) has not documented its own search, and its "we reviewed the literature" is unauditable [15][16].
Fig 1. PRISMA 2020 flow diagram template – the four-box audit trail (identification → screening → eligibility → included) that a compliant systematic review must publish filled in. (Manchikanti, 2024, p. 134) [MEDLIB]
> Fuentes: Manchikanti, Essentials of Interventional Techniques in Managing Chronic Pain (2024) [57]
Fig 1 is the empty template: when you read an SR, mentally fill each box; the reviews that cannot be filled are the ones to distrust [15].
Product class decides the evidence bar, and it is the single most misread point of the frame [A]. In the EU, the same aesthetic clinic uses two legally distinct product classes with opposite evidence regimes:
- Botulinum toxin is a medicinal product, authorised through the medicines route (AEMPS in Spain, EMA centrally). Its evidence bar is a registration dossier of RCTs against its authorised indication; using it in an unlisted facial area is off-label and pulls in RD 1015/2009 [30][31].
- Dermal fillers, PDO/PLLA threads and most energy devices are medical devices under EU MDR 2017/745. Since MDR replaced the old directives, a device needs a clinical evaluation and CE mark, but historically that bar was far lower than a medicines dossier: a filler could be CE-marked on biocompatibility and equivalence data, not on RCT efficacy [28]. This is why the filler literature is thin: the law never required an efficacy RCT to put the product on the market.
Reading an aesthetic RCT against CONSORT: the nine questions, in order [A]. Most aesthetic trials fail at question 2 or 3 [14][9]:
- Who paid and who wrote it? First, not last (see B9.5).
- Was the primary outcome pre-specified and registered? Compare the registry entry (ClinicalTrials.gov / EudraCT / CTIS) to the paper. No registration is itself a finding [27][45].
- What scale measured the result, and what is that scale worth? GAIS (5-point, ceiling-low, almost always positive), validated photonumeric scales (0–4, validated per region and per manufacturer, not interchangeable) [13], patient satisfaction (expectation-sensitive), FACE-Q/BODY-Q (the best available PRO, psychometrically validated, rarely used by industry) [12], objective measures (3D volumetry, cutaneous biometry, ultrasound). Always ask: what does one point on this scale mean for the patient in front of me? [39]
- Who evaluated? An independent, blinded rater on randomised photographs is the standard; the injector grading his own result measured expectation [14].
- How was it blinded? Patient-blinding is nearly impossible (they feel the injection, see the volume); injector-blinding is impossible across two products with different syringes. Mitigations: split-face (each hemiface a product, the best design because each patient is his own control, but confounded by systemic/diffusion effects, critical for toxin) [51]; vehicle/placebo control (ethical, underused); delayed control (the control group treated at the end of follow-up).
- How many patients, and where did the number come from? No sample-size calculation, no study; n=10–20 detects only huge effects [9].
- How long was follow-up? The quietest structural bias in the field: mean follow-up too short for the two things that matter most, real duration of effect and late complications (PLLA nodules, granuloma at 6–24 months). When you read "well tolerated", look at the follow-up first [10].
- How many dropped out, and by what analysis? Losses >20% threaten the result; per-protocol overstates effect because the dissatisfied leave; demand intention-to-treat [44].
- Who does this population resemble? Phototype, age, sex, baseline severity and injector. A result from three 20-year-veteran KOLs does not transfer to your chair; external validity is the weakest link.
For SRs, add two checks: is there a funnel plot or formal publication-bias test [41], and is I² so high that pooling is meaningless [40]? Pooling studies with different products, doses and planes yields a number that describes nothing.
What each reporting guideline actually demands, the high-yield items [A]:
- CONSORT 2010 [14]: a flow diagram of enrolment/allocation/follow-up/analysis; the pre-specified primary outcome; the randomisation and allocation-concealment method; who was blinded; the estimated effect with its CI (not just a p value); and harms. Item 17 requires results for each group, which stops the single-arm "improvement" claim.
- SPIRIT 2013 [19]: the protocol counterpart of CONSORT; a compliant trial has a registered, dated protocol before enrolment, so outcome-switching can be detected by comparison.
- PRISMA 2020 [15][16]: a reproducible search (databases, dates, full strings), the flow diagram of Fig 1, the risk-of-bias method, and a statement of certainty (usually GRADE). A "narrative review" that skips these is opinion with references.
- STROBE [17]: for cohort/case-control/cross-sectional, forces declaration of confounders, how they were handled, missing data, and sensitivity analyses.
- STARD 2015 [18]: for diagnostic accuracy, the reference standard, the patient spectrum, and how indeterminate results were handled, all of which inflate accuracy if hidden.
- CONSORT-Harms [20]: adverse events named, defined, and counted per arm, the single most-omitted section in aesthetic trials given the short follow-up.
ICH-GCP and Helsinki are the ethics floor under all of it [A]. The Declaration of Helsinki (2013 revision) sets the principles (independent ethics review, informed consent, favourable risk/benefit, registration and publication of results) and ICH-GCP E6 operationalises them for trials of medicines; a study that cannot show ethics-committee approval and consent is not merely poorly reported, it is not appraisable, because its data provenance is unknown [26]. Prospective registration is the enforcement lever: the ICMJE will not publish an unregistered trial, which is what makes outcome-switching detectable at all [27].
Worked CONSORT read of a typical filler RCT (P): a "24-week, split-face, rater-blinded RCT of filler A vs filler B, n=30, sponsored by A's maker" fails predictably at: item 6 (primary outcome is GAIS at week 24, a ceiling-low scale, with no MCID stated); item 7 (no sample-size calculation, so n=30 detects only a large difference); item 13 (four dropouts analysed per-protocol, not ITT); item 22 (conclusion favours A on a secondary timepoint while the primary was non-significant). None of these is fraud; each is a reporting choice the guideline exists to expose [14][44][45].
Product class, concretely [A]: botulinum toxin type A products are medicines (medicines route, AEMPS/EMA, authorised indications such as glabellar lines); hyaluronic-acid and calcium-hydroxylapatite fillers, PDO/PLLA threads, and most laser/RF/HIFU/ultrasound devices are medical devices under MDR 2017/745 [28][30]. The practical asymmetry: the toxin's glabellar indication rests on registration RCTs, while the filler beside it on the same tray may hold a CE mark earned on biocompatibility and equivalence, with no efficacy RCT ever required. When you appraise, ask first which legal class the product is, because it tells you what evidence must legally exist [28][30].
Finding the right norm: the EQUATOR Network is the lookup [A]. Rather than memorise the checklists, use EQUATOR as the index: it hosts every reporting guideline and tells you which applies to the design in front of you (CONSORT for an RCT, PRISMA for an SR, STROBE for observational, STARD for diagnostic, SPIRIT for a protocol) [33]. The practical move: identify the design first, pull the matching guideline, and read the paper against its checklist; a study that ignores its own applicable guideline has told you something before you reach the results.
Reading an observational or diagnostic aesthetic paper against its norm [A]: for a cohort of filler outcomes, STROBE forces the questions the abstract hides (how were confounders like injector and baseline severity handled, how much data was missing, was there a sensitivity analysis) [17]; for an ultrasound-localisation accuracy study, STARD forces the reference standard, the patient spectrum and the handling of indeterminate scans, each of which inflates apparent accuracy if concealed [18]. The norm is not bureaucracy: it is the list of exactly the places a weak study goes quiet.
Reporting norm vs legal norm, kept distinct (P): a study can be legally impeccable (ethics approval, CE-marked product, registered trial) and still be a poor report (no CONSORT flow, switched outcome, spun conclusion), and vice versa. The appraiser holds the two axes apart: the legal axis tells you what evidence must exist and whether the product may be used; the reporting axis tells you whether this study is trustworthy. Confusing them is how "it is CE-marked" gets mistaken for "it works" and how "the trial was small" gets mistaken for "the product is illegal" [14][28].
Classic trap: treating a CE mark as evidence of efficacy. CE marks device conformity (safety, biocompatibility, performance claims), not clinical superiority; a CE-marked filler is legal to sell, which says nothing about whether it outperforms the next one [28].
B9.3 · The step-by-step procedure
The procedure is Sackett's 5A cycle, one step per row, with the concrete tool and the abort condition for each [1]:
| Step | Do | Tool | Abort / red flag |
|---|---|---|---|
| 1 Ask | Turn the doubt into a PICO question | PICO worksheet (B9.4) | A vague question returns vague literature |
| 2 Acquire | Descend the source hierarchy until one answers | IFU/CIMA → guideline → SR → primary | An over-built query returns zero, read as "no literature" |
| 3 Appraise | Rate risk of bias with the design-matched tool | RoB 2 / ROBINS-I / AMSTAR-2 / AGREE II / CASP | No design named → cannot pick the tool |
| 4 Apply | Transfer to this patient, or decide not to | Stable-vs-changing test; off-label rule | Population/injector mismatch = do not transfer |
| 5 Assess | Audit your own outcomes and search habit | Complication log; dated change note | If you never look back, your series lies to you |
Step 1 · Ask (PICO). Frame every clinical doubt as Population, Intervention, Comparator, Outcome before typing. "Does PLLA work?" is unsearchable; "In a woman of 55 with midface volume loss (P), does poly-L-lactic acid (I) versus calcium hydroxylapatite (C) give greater rater-blinded improvement at 12 months with fewer late nodules (O)?" names its own answer set. Question type picks the ideal design: therapy → RCT/SR; harm → cohort/case-control; prognosis → cohort; diagnosis → cross-sectional against a reference standard [2].
Step 2 · Acquire: the consultation hierarchy, descend one rung only when the last does not answer [A]:
| Order | Source | Use for |
|---|---|---|
| 1 | Technical sheet / IFU (AEMPS-CIMA for medicines; manufacturer IFU for devices) | Dose, indication, contraindication, storage. Always first; also the medicolegal reference [30] |
| 2 | Guidelines / consensus with a year | Complication management, protocols, prevention [25] |
| 3 | Systematic reviews / Cochrane / Epistemonikos | "Does this actually work?" [15] |
| 4 | Primary literature (PubMed/MEDLINE, TRIP) | When the above is silent or dated [9] |
| 5 | Reference monograph | Anatomy, technique, context. Stable over time [10] |
| 6 | Course, congress, opinion ([D]) |
Where to look next. Never closes a question |
Search mechanics, the minimum effective set: start simple ([Author] AND), because over-built queries return zero and a zero from over-building reads as "no literature exists", the commonest way to conclude wrongly; add filters (systematic review, randomized controlled trial, last 5 years); use PubMed Similar articles, the most under-used tool there; save email alerts for your 5–10 core topics, the whole mechanics of "staying current". Full text, not the abstract: the abstract lies by omission, because methods, real follow-up and dropouts live in the body. The open-access chain Unpaywall → OpenAlex → Europe PMC → PMC clears most paywalls; a paywall is an obstacle, not an endpoint.
Step 3 · Appraise: pick the instrument by design, then work its signalling questions [A]. Using the wrong tool (RoB 2 on a cohort, AMSTAR-2 on a single RCT) produces a confident wrong grade:
| Design in front of you | Instrument | Structure | Ref |
|---|---|---|---|
| Randomised trial | RoB 2 | 5 domains, signalling questions → low / some concerns / high | [21] |
| (older RCT appraisal) | Cochrane RoB 1 | 7 domains, judged high/low/unclear | [24] |
| Non-randomised intervention | ROBINS-I | 7 domains, confounding first, vs a target trial | [22] |
| Systematic review | AMSTAR-2 | 16 items, 7 critical; one critical flaw → low confidence | [23] |
| Clinical guideline | AGREE II | 23 items, 6 domains, two appraisers | [25] |
| Diagnostic-accuracy study | STARD 2015 / QUADAS-2 | reference standard, spectrum, flow | [18] |
| Any design, quick | CASP checklists | 3 questions: valid? what are the results? applicable? | [34] |
Worked order for an RCT with RoB 2 [21]: (D1) randomisation process, was allocation concealed and were baseline groups balanced; (D2) deviations from intended intervention, was analysis intention-to-treat; (D3) missing outcome data, is the result robust to the losses; (D4) measurement of the outcome, was the assessor blinded and the method the same in both arms; (D5) selection of the reported result, does the paper match its registered protocol. Any single domain "high" makes the overall judgement "high risk". For a systematic review with AMSTAR-2, the seven critical domains (protocol registered a priori, adequacy of the search, justification of exclusions, risk-of-bias assessment, meta-analytic method, risk of bias considered when interpreting, publication bias assessed) govern the verdict: one critical flaw drops the whole review to "critically low confidence", regardless of the other nine items [23].
Step 4 · Apply. Run the stable-vs-changing test (B9.1) to set the currency bar, check the technical sheet for indication, and if the claim is off-label decide explicitly under RD 1015/2009 with documented justification and specific consent, never by routine [31]. Transfer only when population, severity and injector match; otherwise record why you did not.
Step 5 · Assess. Audit your own complication rate and outcomes on a schedule (B9.6). And the discipline that closes the loop: when you change something in your practice, write why and with what source; in two years you will not recall whether you read it or were told it, and that difference is what this chapter preserves.
The 11-step protocol for a new claim (a rep, a colleague or a lecture says X works) [MODELO]: (1) stable or changing claim? set the currency bar. (2) who says it, what do they gain? conflicts first. (3) find the IFU/technical sheet; in indication or off-label (RD 1015/2009)? [31] (4) find a dated guideline/consensus. (5) find an SR. (6) if nothing above, go to primary and run the nine questions of B9.2. (7) look at follow-up before the result. (8) translate the effect to chairside language: how much improvement, in how many patients, for how long, at what risk, at what cost. (9) assign an [A]–[D] tag and write it beside the claim. (10) if evidence is low but you act anyway, make it an explicit, documented decision, not an omission. (11) tell the patient in those terms: "the evidence is limited and here is what we know" is a sentence you can say in consultation, and it protects you.
Search mechanics, one level deeper [A]. Build the PubMed query from the PICO, not from a sentence:
- Field tags narrow noise:
botulinum toxin[Title/Abstract] AND glabella*[Title/Abstract] AND randomized[Publication Type]. The truncation*catches glabella/glabellar. - MeSH captures synonyms you would miss: the MeSH term
Cosmetic Techniquesindexes papers that never use your keyword. - Boolean order:
ORwithin a concept (grouped in parentheses),ANDbetween concepts:(filler OR "dermal filler" OR hyaluronic acid) AND (tear trough OR infraorbital). - Pre-appraised layers first: Cochrane Library and Epistemonikos for SRs; TRIP and DynaMed/UpToDate for pre-digested answers; ClinicalTrials.gov / EU-CTIS to find the registered protocol and any unpublished result.
- The commonest failure is the over-built query returning zero, read as "no evidence exists". When you hit zero, remove terms one at a time; most "gaps" are search failures, not literature gaps.
Open-access, the escalation that beats a paywall [A]: a locked DOI is not the end. Run Unpaywall → OpenAlex → Europe PMC → PMC, then the author's institutional repository, then a direct request to the corresponding author. The abstract is never a source: the short follow-up, the per-protocol switch and the dropout count live in the methods and results, not the summary.
Fully worked RoB 2 appraisal (P) [21], applied to the split-face filler RCT of B9.2:
- D1 Randomisation: "randomised by coin flip at the visit" → allocation not concealed (the injector knew the next assignment); baseline unbalanced on severity → some concerns.
- D2 Deviations: three patients received touch-ups off-protocol, analysed as originally assigned only in a footnote → some concerns.
- D3 Missing data: four of thirty dropped out, all from arm B, analysed per-protocol → the result is not robust to the losses → high.
- D4 Outcome measurement: rater blinded to arm but the two fillers produce visibly different swelling, partially unblinding by proxy → some concerns.
- D5 Selective reporting: registry lists GAIS at week 24 as primary; paper headlines patient satisfaction at week 12 → high.
- Overall: the worst single domain governs → high risk of bias. The verdict is reproducible because it names where (D3, D5), which "it looks weak" never does.
For a non-randomised comparison, ROBINS-I reframes the whole appraisal around a target trial: describe the ideal RCT you would have run, then judge how far the observed study departs, with confounding as the first and heaviest domain (did the two groups differ in baseline severity, injector skill or expectation?) [22]. This is also the correct lens for real-world registry data (B9.8), which its sample size tempts a reader to over-trust.
Acquire smarter: the pre-appraised pyramid (6S) [A]. Effort is saved by starting at the top, where someone has already appraised and synthesised:
| Layer (top = least work) | Example | Aesthetic reality |
|---|---|---|
| Systems | Decision support wired to the record | Effectively absent in aesthetics |
| Summaries | UpToDate / DynaMed topic | Sparse for injectables; good for adjacent derm |
| Synopses of syntheses | DARE-style SR abstracts, ACP Journal Club | Rare |
| Syntheses | Cochrane / Epistemonikos SR | Few, but the ceiling of the field |
| Synopses of studies | Structured abstracts | Common, low value alone |
| Studies | The primary RCT/series | Where most aesthetic evidence lives, so appraise it yourself |
The lesson: in aesthetics you are forced further down the pyramid than in most of medicine, so the appraisal skills of B9.2 and Step 3 are not optional; there is rarely a summary to trust above you.
Match the question to the design to the tool, in one line [A] [2][21][22][42]: therapy → RCT/SR → RoB 2/AMSTAR-2; harm → cohort/case-control (or RCT harms) → ROBINS-I; prognosis → cohort → ROBINS-I with confounding foremost; diagnosis → cross-sectional vs reference standard → STARD/QUADAS-2. Picking the tool is not bureaucracy: RoB 2 asks about randomisation, which a cohort has none of, so using it on a cohort produces a meaningless "high risk" that hides the real question (confounding) that ROBINS-I would have asked.
Staying current is a mechanic, not a virtue [A]. You cannot read the field, and you do not need to; you need a system that pings you when something you do daily has changed. Save PubMed email alerts for your 5–10 core queries, subscribe to the tables of contents of the two journals that publish your subfield, and set a "last updated" watch on any living guideline you follow [52]. The whole of "keeping up" is those alerts plus the quarterly CIMA re-check; without the automation it does not survive a clinic week [30].
The whole cycle on one claim (P): rep says "biostimulator X gives 18-month results". Ask: in my patients, does X versus my current product give longer rater-blinded improvement at 18 months? Acquire: IFU (indication, dose), then Cochrane/Epistemonikos, then PubMed; the 18-month claim traces to one open-label series. Appraise: ROBINS-I → high risk (no control, confounding, self-assessed). Apply: changing claim, currency bar high, evidence low; if I use it, off-label decision documented under RD 1015/2009. Assess: log outcomes with a denominator and re-check in a year. Tag: [D] slide upgraded to [B] low-certainty series, never [A]. That is the chapter, executed.
Classic trap: appraising with no instrument, or the wrong one. "It looks like a good study" is not appraisal; the point of RoB 2 / ROBINS-I / AMSTAR-2 is to make the judgement reproducible by naming the domain where the study fails [21][23].
B9.4 · Templates and documents
The instruments of B9.3 turned into fill-in forms. Each is copy-pasteable into a note; all reference values sit in B9.6.
Template 1 · PICO worksheet [MODELO] (drives the search):
P (patient/problem): age, sex, phototype, region, severity, prior treatment
I (intervention): product/device, dose/energy, plane, technique
C (comparator): placebo / another product / no treatment / standard of care
O (outcome): measure + timepoint + who rates it (blinded?) + MCID
Question type: [ ] therapy [ ] harm [ ] prognosis [ ] diagnosis
Ideal design: therapy→RCT/SR · harm→cohort · diagnosis→cross-sectional
Template 2 · CASP-style RCT appraisal one-pager [A] [34][14]:
A. Is the basic design valid?
1. Focused question (PICO)? Y / N / can't tell
2. Randomised, allocation concealed? Y / N / ?
3. All entrants accounted for at end (ITT)? Y / N / ?
B. Was bias minimised?
4. Patients/clinicians/assessors blinded (which)? Y / N / ?
5. Groups similar at baseline? Y / N / ?
6. Groups treated equally apart from the intervention? Y / N / ?
C. What are the results?
7. Effect size (absolute, not only relative)? ______
8. Precision (95% CI width; crosses null?)? ______
9. Primary outcome pre-registered = published? Y / N / ?
D. Will they help my patient?
10. Population/injector applicable to my chair? Y / N / ?
11. Benefits vs harms vs cost acceptable? Y / N / ?
Overall risk of bias: LOW / SOME CONCERNS / HIGH
Template 3 · RoB 2 signalling grid [A] [21] (RCT):
| Domain | Signalling question (short) | Judgement |
|---|---|---|
| D1 Randomisation | allocation concealed, baseline balanced? | Low / Some / High |
| D2 Deviations | analysed by ITT, no differential co-intervention? | Low / Some / High |
| D3 Missing data | outcome available for ~all; robust to losses? | Low / Some / High |
| D4 Outcome measurement | assessor blinded, method identical in arms? | Low / Some / High |
| D5 Selective reporting | result matches registered protocol? | Low / Some / High |
| Overall | worst single domain governs | Low / Some / High |
For non-randomised comparisons, swap in ROBINS-I: its first and heaviest domain is confounding, judged against the hypothetical target trial you would have run [22]. For a systematic review, use AMSTAR-2 and record which of the 7 critical items failed [23].
Template 4 · GRADE Summary-of-Findings (SoF) row [A] [3][6]. One row per outcome; this is how a guideline states what it knows:
Outcome: __________________ Timepoint: ______
Absolute effect: control ___ per 1000 → intervention ___ per 1000 (Δ ___)
Relative effect: RR/OR ___ (95% CI ___ to ___)
Nº participants (studies): ___ ( ___ RCTs / ___ observational )
Certainty (GRADE): ⊕⊕⊕⊕ high · ⊕⊕⊕◯ moderate · ⊕⊕◯◯ low · ⊕◯◯◯ very low
Downgraded for: [ ]RoB [ ]inconsistency [ ]indirectness [ ]imprecision [ ]pub.bias
Plain-language summary: "________ probably/may improve/little-or-no difference"
The Evidence-to-Decision (EtD) frame then adds, on top of the SoF: benefit/harm balance, values, resources, equity, acceptability, feasibility, and yields strong or conditional [7][8].
Template 5 · 2×2 diagnostic table [A] [42][43] (fill from any diagnostic-accuracy study; formulas in B9.6):
Disease + Disease −
Test + TP FP → PPV = TP/(TP+FP)
Test − FN TN → NPV = TN/(FN+TN)
Sens=TP/(TP+FN) Spec=TN/(TN+FP)
LR+ = Sens/(1−Spec) LR− = (1−Sens)/Spec (interpret in B9.6)
Template 6 · Forest-plot reading checklist [A] [40][41] (for any meta-analysis figure):
[ ] Vertical line of no effect at 1 (ratio) or 0 (difference) identified
[ ] Each square = one study; square area = weight (larger = more weight)
[ ] Horizontal line = that study's 95% CI (does it cross the null?)
[ ] Diamond = pooled estimate; its width = pooled 95% CI
[ ] I² and its band read (see B9.6); >50% → pooling questionable
[ ] Funnel plot / Egger test present for publication bias?
[ ] Are the pooled studies clinically poolable (same product/dose/plane)?
Template 7 · Conflict-of-interest disclosure (ICMJE-style) [A] [26][47], to file for any talk or publication you give and to look for in any you read:
Funding source of this work: __________ Role of funder in design/analysis? __
Author payments (last 36 mo): speaker fees / advisory board / equity / travel
/ free product / research grant (per company)
Data analysed by: independent statistician / sponsor (circle)
Trial registration nº: __________ Protocol matches report? Y / N
Template 8 · Dated practice-change note [MODELO] (the memory device of Step 5):
Date: ____ Change: ____ Trigger: (paper/guideline/course) ____
Source + tag [A–D]: ____ Evidence certainty: high/mod/low/very-low
Reviewed on: ____ (re-check technical sheet quarterly, CIMA)
Template 9 · PRISMA search log [A] [15] (so your own "I reviewed the literature" is auditable):
Databases + date searched: PubMed __ / Cochrane __ / Epistemonikos __
Full search string: ______________________________
Records found: ___ after de-dup: ___ screened: ___ excluded (reason): ___
Full text assessed: ___ included: ___ Risk-of-bias tool used: ______
Template 10 · GRADE evidence profile (per outcome) [A] [6], the auditable version of Template 4:
Outcome | Nº studies (design) | RoB | Inconsistency | Indirectness |
| Imprecision | Publication bias | Upgrades | → Certainty (⊕)
Each column: "not serious / serious / very serious"; each downgrade −1 or −2.
Observational may upgrade for: large effect / dose-response / opposing confounding.
Template 11 · QUADAS-2 / STARD diagnostic checklist [A] [18] (for a diagnostic-accuracy paper, e.g. ultrasound for filler localisation):
[ ] Reference standard defined and independent of the index test?
[ ] Consecutive/representative patient spectrum (not case-control of extremes)?
[ ] Index test interpreted blind to the reference standard?
[ ] All patients got the same reference standard? Indeterminates reported?
[ ] 2×2 table reconstructable (TP/FP/FN/TN)? → compute LR+/LR− (B9.6)
Template 12 · AGREE II guideline read [A] [25] (score each domain 1–7, two appraisers):
D1 Scope & purpose | D2 Stakeholder involvement | D3 Rigour of development |
D4 Clarity | D5 Applicability | D6 Editorial independence
Red flags: panel single-brand? funding undisclosed? no systematic evidence base?
Overall: recommend / recommend with modifications / do not recommend
Template 13 · Shared-decision script for low evidence [MODELO] (the sentence that informs and protects):
"The evidence for [procedure] is [level A–D, certainty ___]. On average it
[absolute benefit, in patients-per-100], lasts about [___], with a [___%] risk
of [___]. What we do not know is [orphan question]. Given that, do you want to
proceed?" → documented in the record with source + tag.
Template 14 · Practice audit / complication log [MODELO] (the Assess metric with a denominator):
Procedure | Nº done (denominator) | Complications (type, nº) | Rate =
Revision/retreatment nº | PRO (FACE-Q) baseline→follow-up | Reviewed on: ___
Without the denominator, a complication count is a number with no meaning; the rate is what compares to the literature.
Template 15 · Level-of-evidence quick-assign card [A] [2][6], to stamp a claim in seconds at the point of care:
Design in front of me: ___________________
OCEBM level (shorthand): 1 SR-RCT / 2 RCT / 3 cohort / 4 series / 5 mechanism
GRADE certainty (for a recommendation): high / mod / low / very low
Any downgrade? RoB / inconsistency / indirectness / imprecision / pub-bias
Atlas tag written beside the claim: [A] / [B] / [C] / [D] Year: ____
Template 16 · PICO-to-search-string converter [A] [15] (worked, so the search is reproducible):
PICO: P= midface volume loss, 50s woman | I= PLLA | C= CaHA | O= rater-blinded improvement 12 mo
String: ("poly-L-lactic acid" OR PLLA OR Sculptra) AND (calcium hydroxylapatite
OR CaHA OR Radiesse) AND (midface OR cheek) AND (randomized OR comparative)
Filters: last 5 y · RCT/SR · humans Databases: PubMed + Cochrane + Epistemonikos
Template 17 · Currency triage card [MODELO] (which claims decay, from B9.1):
STABLE (a dated old source still serves): anatomy · mechanism · basic pharmacology
CHANGING (needs a current source): regulation · availability · IFU/dose · device
parameters · complication protocols
Rule: a 2009 source is fine for a STABLE claim if the year is visible; never for CHANGING.
Template 18 · N-of-1 protocol skeleton [A] [54] (for a chronic, reversible, self-rated intervention):
Patient + question (PICO): ___________________
Treatment vs control periods: ≥3 pairs, order randomised, washout between
Blinding: patient/rater blinded where feasible (identical vehicle)
Outcome: patient-rated on a validated scale, same instrument each period
Analysis: within-patient comparison; decision rule agreed in advance
Template 19 · Registration-vs-report comparison [A] [27][45] (the single highest-yield anti-spin check):
Registry ID (ClinicalTrials.gov / EudraCT / CTIS): __________
Registered PRIMARY outcome + timepoint: __________
Published PRIMARY outcome + timepoint: __________
Match? Y / N (N = outcome switching → distrust the headline)
Registered n / analysed n: ___ / ___ Any added post-hoc outcomes? Y / N
Template 20 · One-page appraisal verdict [MODELO] (what you file after reading a paper):
Citation + tag [A–D]: __________ Design + OCEBM level: __________
Risk of bias (tool + verdict): RoB2/ROBINS-I/AMSTAR-2 → low/some/high
Effect (absolute + CI): __________ MCID met? Y / N
Conflicts / funding: __________ Registration matches? Y / N
Applies to my patient? Y / N GRADE certainty for use: high/mod/low/v-low
One-line verdict: __________
Worked Template 20, filled (P) (the split-face filler RCT of B9.2–B9.3, appraised): design = split-face RCT, OCEBM level 2; risk of bias = RoB 2 → high (D3 per-protocol on differential dropout, D5 outcome switching); effect = a GAIS difference with a wide CI and no stated MCID; conflicts = maker-funded, analysis by sponsor; registration = primary switched from week 24 to week 12; applies to my patient = partly (expert injectors, small n); GRADE certainty for use = low; one-line verdict = "signal of benefit, high risk of bias and spin, treat as hypothesis-generating, not practice-changing". The value of the form is that the verdict names why, so a colleague can check it.
The templates above are the chapter's operational core: B9.3 is the procedure, these are the forms that make each step leave a record instead of a memory.
Tools that build these forms for you [A]: GRADEpro produces the SoF and EtD tables of Templates 4 and 10; Covidence and RevMan run the PRISMA screening and forest plots of Templates 6 and 9; RoBVis visualises the RoB 2 grid of Template 3. You do not need them to appraise a single paper, but for building or reading a systematic review they turn the templates above into an auditable trail rather than a private judgement [15][21].
How to use these without drowning [MODELO]: not every paper earns all twenty forms. Triage first with Template 15 (level and tag in seconds); if the claim will change what you do, escalate to Template 2 (CASP) or Template 3 (RoB 2) and file Template 20 (verdict). The registration check (Template 19) and the absolute-effect line (Template 2, row 7) are the two you never skip, because they are the two a marketing summary is built to hide.
Classic trap: downloading a template and never filling the absolute effect and registration lines, the two fields a marketing summary always omits. A CASP sheet with rows 7 and 9 blank has appraised nothing [37][45].
B9.5 · Frequent errors and their cost
The error grid [A] (error → why it happens → what it costs → the fix):
| Error | Why it happens | Cost | Fix | Ref |
|---|---|---|---|---|
| Citing a PMID/dose from memory | You recall reading it | Fabricated authority; wrong dose | Verify the identifier before writing it | [1] |
p<0.05 read as "works well" |
Trained reflex | Adopting a trivial effect | Read effect size and 95% CI | [35][36] |
| Relative risk without absolute | It is what the maker gives | Overstated benefit sold to patient | Demand ARR; "50%" can be 0.02%→0.01% | [38] |
| Reading only the abstract | Time and paywalls | Missing the short follow-up and dropouts | Full text via Unpaywall/PMC | [9] |
| "Well tolerated" without follow-up | Sounds like a conclusion | Missing a late nodule/granuloma | Read follow-up before the safety line | [10] |
| Comparing products across studies | Marketing tables do it | Non-comparable scales/populations | Only head-to-head trials compare | [40] |
| Split-face treated as two samples | Doubles the apparent n | False significance | Paired analysis, per patient | [51] |
| Ignoring the disclosure | Small print at the end | Reading a marketing study as science | Read conflicts first | [47] |
| Taking a consensus as evidence | It has guideline format | Inheriting the panel's conflicts whole | Check panel composition and funding | [25] |
| "Not in the literature" after one search | Over-built query returned zero | A real answer missed | Simplify the query first | [1] |
| Using your own series as proof | It is what you see | Survivorship bias baked in | The dissatisfied do not return | [11] |
| Adopting a just-learned technique | Post-course enthusiasm | A fad on the first patient | Search its evidence before patient one | [9] |
| Scale validated for another region | It is to hand | Uninterpretable result | Photonumeric scales validate per region/product | [13] |
| Multiple comparisons uncorrected | Many endpoints reported | Chance "findings" | 60 tests → ~3 significant by luck | [49] |
| Before-after without a control | It is easy to run | Regression to the mean sold as effect | Demand a control arm | [50] |
Industry influence: the structural fact, quantified [A]. Almost all aesthetic research is funded by the seller of the product; there is no meaningful public funding to compare two fillers. This is not an insinuation, it is the field's funding model, and the effect is measured: industry-sponsored studies more often reach favorable efficacy results, RR 1.27 (95% CI 1.17–1.37), and favorable conclusions, RR 1.34 (95% CI 1.19–1.51), than non-industry studies; the same reviews found industry studies more often had low risk of bias from blinding (RR 1.25) yet less agreement between their results and their conclusions (RR 0.83): the spin is in the abstract, not the data table [47]. Crucially, this bias is not caught by standard risk-of-bias tools, so a trial can be "low risk of bias" on RoB 2 and still tilt [47].
Fig 2. The Sugar Industry paper (Kearns 2016) as taught on a UPO deck: sponsored research set to cast doubt on sucrose harm while promoting fat as the culprit, funding undisclosed. (UPO Sorted, Prof. Ayala deck, p. 31) [D]
> Fuentes: 04 PRESENTACION Clase3 Nutricion Antienvejecimiento-Prof Ayala · original: Kearns CE, Schmidt LA, Glantz SA. JAMA Intern Med 2016 [48]
Fig 2 is the field's cautionary archetype [48]: the historical record of the Sugar Research Foundation shaping the coronary-heart-disease literature is the mechanism that plays out, quietly, wherever one funder owns most of a field's trials. In aesthetics the funder is often also the educator: the course where you learned the technique was paid for by the seller of the product of the technique.
The six mechanisms, most to least visible [A] [45][46][47]:
| Mechanism | How it shows |
|---|---|
| Publication bias | What fails is not published; three positive trials can hide seven negatives in a drawer [41] |
| Chosen comparator | Against placebo or a sub-optimal rival dose, never the best competitor at optimal dose |
| Chosen outcome / spin | A scale sensitive to the sponsor's effect; a negative primary spun into a positive secondary [45][46] |
| Ghost/guest authorship | Written by a medical-communications agency, signed by the KOL; acknowledgements sometimes betray it |
| KOL incentive chain | Speaker fees, advisory boards, free product, travel, training sponsorship |
| Training as sales channel | The maker is the teacher; technique and brand are learned inseparably |
What this does NOT mean (P): funded research is not automatically false. It is often better resourced and more closely monitored than independent work, and a conflicted KOL is not lying. It means the set of questions asked is skewed, and the questions no maker wants answered simply have no literature: how much cumulative product is too much, what happens at 15 years, does treating less give a better long-term result, is the cheap product worse. These are the field's orphan questions, and their emptiness is a finding, not a gap in your search [11][47].
The statistics traps, in detail [A]:
- Statistical vs clinical significance. The number-one error. p<0.05 says only that the data are improbable under the null; with a large sample, an irrelevant difference is "significant" [35].
- What a p value actually is. The probability of a result at least this extreme if the null were true. Not the probability the treatment works, nor that the result is false; p=0.04 and p=0.06 are not different worlds [35].
- Relative vs absolute risk, the commonest inflation. "Reduces risk 50%" can mean 0.02%→0.01%. Demand the absolute; without it the relative figure persuades but does not inform [38].
- Multiple comparisons without correction. Five regions × four timepoints × three scales = 60 tests; by chance about three come out significant. A paper that runs many and highlights two is fishing [49].
- Before-after without a control. The patient improves by regression to the mean, by the photograph, by the care received and by wanting it to work; with no control arm, the study measures the sum of all of that [50].
- Split-face treated as independent. Two hemifaces of one patient are not two patients; counting them as such doubles the apparent n and inflates significance. Use paired analysis [51].
- Surrogate and short-horizon endpoints. A 12-week wrinkle-score change is a surrogate for the patient's real question (how long, how safe over years); switching from ITT to per-protocol quietly overstates effect [44].
Six more biases that specifically distort aesthetic reading [A]:
| Bias | Mechanism | Aesthetic example |
|---|---|---|
| Survivorship | Only successes remain visible | Your clinic photo wall; the dissatisfied left and did not return [11] |
| Immortal-time | A period where the outcome could not occur is misassigned | "Patients who completed 3 sessions did better" (dropouts could not, by definition) |
| Simpson's paradox | A pooled trend reverses within subgroups | A filler "better overall" is worse in every phototype once stratified |
| Ecological fallacy | Group-level correlation read as individual | "Countries using product X have smoother skin" says nothing about a patient |
| Base-rate neglect | Ignoring prevalence when reading a test/claim | A "positive" screening for a rare complication is mostly false positives [42] |
| Regression to the mean | Extreme baselines drift toward average | The worst wrinkles "improve" without any treatment [50] |
Spin, quantified [A]: Boutron showed that trials with a non-significant primary outcome routinely present the result as if it were positive (spin in the abstract conclusion), and readers shown the spun version rate the treatment more favourably [46]. Selective outcome reporting is its structural twin: comparing published papers to their protocols, primary outcomes are silently changed in a large share of trials, almost always toward the significant result [45]. This is why Template 2, row 9 (registered primary = published primary) is the highest-yield single check in an appraisal.
The cost, made concrete (P). The errors above are not academic: each has a price. A relative-risk read as absolute sells a patient a benefit that is not there (informed-consent failure, medicolegal exposure). "Well tolerated" on a 12-week study puts a late-nodule risk on a patient who was told there was none. Adopting a just-learned course technique before searching its evidence is how a fad reaches the first patient at their expense and yours. And using your own series as proof, unaware it is survivorship-biased, entrenches a technique that the returning-patient sample can never falsify [11]. The economic cost compounds the clinical one: a product bought on a relative-risk claim, a device bought on a sponsored trial, a training paid to the seller. Reading past marketing is not cynicism; it is the fiduciary duty the consultation carries.
Reading conflicts in practice, the five checks [A] [47]: (1) disclosure first, before the abstract; its absence in a product study is itself a datum. (2) check the comparator: placebo for an established product is a marketing study. (3) check who analysed the data: "statistical analysis performed by the sponsor" changes the read. (4) find the trial registration and compare outcomes. (5) apply the same rule to guidelines and to this atlas: a consensus of ten same-brand speakers is not a field consensus, and the atlas declares school sponsorship on the first line for exactly this reason.
The aesthetic-marketing tells, a field guide [A] (each maps to a mechanism above): "clinically proven" with no citation (publication/outcome choice); a before/after with changed lighting, angle or expression (no control, regression to the mean) [50]; "50% more collagen" with no absolute number or histology denominator (relative inflation) [38]; "preferred by physicians" (KOL incentive chain); "n=2000 treatments" as if volume were rigour (survivorship, no comparator) [11]; "no downtime, well tolerated" from a 4-week study (short follow-up) [10]; a head-to-head win against a rival at a sub-optimal rival dose (chosen comparator). None is illegal; each is a reading task, and the fix is always the same: find the absolute effect, the comparator, the follow-up and the registration [37][45].
Reappraising your own testimonials (P): the photo wall is the purest survivorship sample you own, because the patients who were unhappy did not return to be photographed, and the ones who did selected themselves for satisfaction. It is evidence of what a good result looks like, never of how often you achieve it; only the audited rate with a denominator (Template 14) answers that [11].
The cost, tallied in one line per error class (P): a statistics error costs a wrong clinical decision; a conflict-of-interest blind spot costs a product bought on marketing; a follow-up blind spot costs a late complication the patient was told could not happen; a survivorship blind spot costs a technique entrenched that the data could never falsify. Each is preventable by one named check from the grid above, which is why the grid is worth more than any single fact in it.
Your own conflicts, handled [A]: the rule that applies to the industry applies to you. Declare conflicts in any talk or publication, refuse remuneration tied to volume prescribed or injected, and separate the evidence from the marketing when you present. The atlas names brands (a private wiki), but keeps the evidence lane and the marketing lane apart on purpose; a clinician who blurs them in their own practice has become the thing B9.5 teaches you to read past [47].
Classic trap: believing that "low risk of bias" clears a trial of industry influence. It does not: Lundh found the sponsorship effect survives a clean risk-of-bias assessment, so the conflict must be judged separately, first [47].
B9.6 · Metrics: what is measured and its reference value
The metric ladder [A] (what it is → reference value / how to read it):
| Metric | What it is | Reference value / reading | Ref |
|---|---|---|---|
| RR (risk ratio) | risk in treated ÷ risk in control | 1 = no effect; <1 = protective | [37] |
| OR (odds ratio) | odds ratio; approximates RR only when the outcome is rare | 1 = no effect; overstates RR when common | [43] |
| HR (hazard ratio) | ratio of event rates over time | 1 = no effect; needs a time axis | [40] |
| ARR (absolute risk reduction) | control risk − treated risk | the honest number; demand it | [38] |
| RRR (relative risk reduction) | ARR ÷ control risk | inflates; "50%" hides the base rate | [38] |
| NNT / NNH | 1 ÷ ARR (or ÷ ARI for harm) | lower NNT better; pair with time + base risk | [37] |
| 95% CI | range compatible with the data | crosses 1 (ratio)/0 (difference) → non-significant | [36] |
| p value | P(data this extreme │ null true) | 0.05 is a convention, not a boundary of truth | [35] |
| Effect size / Cohen's d | standardised magnitude of the shift | ~0.2 small, 0.5 medium, 0.8 large (context-bound) | [39] |
| MCID | smallest change a patient notices/values | if none published, the scale cannot prove usefulness | [39] |
| I² | % of variance from heterogeneity | 0–40 maybe; 30–60 moderate; 50–90 substantial; 75–100 considerable | [40] |
| Sensitivity / specificity | TP/(TP+FN) · TN/(TN+FP) | fixed test properties; do not move with prevalence | [42] |
| PPV / NPV | TP/(TP+FP) · TN/(FN+TN) | move with prevalence; low prevalence sinks PPV | [42] |
| LR+ / LR− | Sens/(1−Spec) · (1−Sens)/Spec | LR+ >10 or LR− <0.1 large; 5–10/0.1–0.2 moderate | [43] |
| κ (kappa) | inter-rater agreement beyond chance | <0.2 poor, 0.4–0.6 moderate, >0.8 very good | [12] |
Risk measures, and the inflation to watch. RR, OR and HR are ratios: the point estimate is meaningless without its CI. The single most manipulated pair is ARR vs RRR: a 50% RRR sounds decisive and can be a move from 0.02% to 0.01%; the ARR (0.01%) and its NNT (1 in 10,000) tell the patient the truth [38][37]. NNH is directly usable for complications and under-used: translate "0.3 % risk of nodule" into "1 in about 333", a sentence a patient understands [37].
Precision beats the p value. The 95% CI is more informative than p and is published less because it exposes uncertainty: a 95% CI and p=0.05 are mathematically equivalent, but the CI is in the units of the outcome, so a clinician reads it directly [36]. Read the extremes: if the CI runs from "trivial" to "enormous", the study is compatible with the treatment barely working. The ASA's formal statement is the anchor: a p value neither measures the probability that the hypothesis is true nor the size or importance of an effect, and "statistical significance" is not scientific or clinical importance [35].
Fig 3. A real forest plot read panel by panel: study squares sized by weight, whiskers as 95% CIs, subgroup and overall diamonds, pooled RR 7.13 (3.92 to 12.95), I² and the Z-test for overall effect [58]. (Clinical Anesthesia, 5th ed., 2006, p. 148) [MEDLIB]
> Fuentes: Clinical Anesthesia, 5th ed. (2006) [58]
Fig 4. The reading key in the abstract: square area = weight, whisker = CI, dotted vertical = no effect (OR 1), diamond = pooled estimate; a study whose whisker crosses 1 is individually non-significant. (Cross, 2008, p. 236) [MEDLIB]
> Fuentes: Cross, Physics, Pharmacology and Physiology for Anaesthetists (2008) [59]
Fig 3 is a real meta-analysis and Fig 4 its schematic key: together they show every element to read: the vertical line of no effect (RR/OR = 1), each square whose area is the study's weight, the horizontal CI whisker (crossing the line = that study is non-significant), and the diamond whose centre is the pooled estimate and whose width is the pooled CI. In Fig 3 the pooled RR is 7.13 with a 95% CI of 3.92 to 12.95: far from 1, precise, and significant [40].
Heterogeneity gates pooling. I² is the share of variability due to real between-study difference rather than chance; the Cochrane bands are 0–40% (might not be important), 30–60% (moderate), 50–90% (substantial), 75–100% (considerable) [40]. In aesthetics, pooling trials with different products, doses and planes produces a high-I² number that describes nothing. Publication bias is screened with a funnel plot and Egger's test: asymmetry suggests small negative trials are missing [41].
Fig 5. Effect size = the standardised distance between the control and treated distributions (μ0 → μ1); overlapping curves with a small shift can still be "significant" with a large n yet clinically invisible. (Schmidt / Bogduk, 2007, p. 740) [MEDLIB]
> Fuentes: Schmidt, Encyclopedia of Pain (2007) [60]
Effect size and MCID. Fig 5 shows why statistical significance and clinical importance diverge: the effect is the displacement of the whole distribution (μ0 → μ1), and with a large sample even a tiny displacement is "significant". The question that rescues the reader is the MCID: the smallest change the patient actually notices or values; a 0.3-point move on a 5-point GAIS can be statistically robust and clinically invisible, and if no MCID has been established for the scale, the study cannot prove usefulness at all [39].
Diagnostic metrics. Sensitivity and specificity are fixed test properties; PPV and NPV move with prevalence, so a test with excellent sensitivity has a low PPV in a low-prevalence setting (most positives are false) [42]. Likelihood ratios combine both and are prevalence-independent: LR+ above 10 or LR− below 0.1 shift probability decisively; 5–10 or 0.1–0.2 moderately; near 1 the test is useless [43]. The reader converts pre-test to post-test probability, which is the whole point of ordering the test.
Fig 6. The descriptive-statistics primer: central tendency vs dispersion, a worked mean and SD, and the stated equivalence of the 95% CI and p=0.05, the base layer under every metric above [61]. (Aitkenhead, 2001, p. 21) [MEDLIB]
> Fuentes: Aitkenhead, Textbook of Anaesthesia 4th ed (2001) [61]
Descriptive base. Fig 6 is the layer under everything else: report a central tendency (mean when symmetric, median when skewed) and a dispersion (SD or IQR); a mean without a spread is misleading, and spurious precision (too many decimals) is a tell of a computer package trusted over judgement.
Metrics of the appraisal itself. RoB 2 yields one of three grades per domain; AMSTAR-2 yields a confidence rating driven by how many of its 7 critical items pass; AGREE II yields a scaled domain percentage from two appraisers [21][23][25]. Metrics of your own practice, the Assess step: complication rate per procedure, revision/retreatment rate, and a validated PRO such as FACE-Q tracked over time; without them your personal series is anecdote [12].
Worked diagnostic example [A] (illustrative cell counts, to show the method of [42][43]): imagine 1000 patients, disease prevalence 10% (100 diseased), a test with sensitivity 90% and specificity 90%. Then TP=90, FN=10, FP=90, TN=810. Sensitivity = 90/100 = 90%; specificity = 810/900 = 90%; PPV = 90/180 = 50% (half of positives are false, because prevalence is low); NPV = 810/820 = 99%. LR+ = 0.9/0.1 = 9 (moderate), LR− = 0.1/0.9 = 0.11 (moderate). The lesson the numbers teach: a "90/90" test still gives a coin-flip PPV when disease is rare, which is why PPV cannot be quoted without the prevalence, and why the LR (prevalence-independent) is the portable number [42][43].
Pre- to post-test probability. The clinical use of a test is to move a probability: pre-test odds × LR = post-test odds. A Fagan-style reading (anchor the pre-test probability, draw through the LR, read the post-test probability) makes explicit that a test with an LR near 1 changes nothing, so ordering it is waste [43]. In aesthetics this governs, for example, whether ultrasound meaningfully changes your estimate that a nodule is product versus abscess before you act.
Worked NNT [A]: if a preventive step drops a complication from 6% to 2%, ARR = 4 percentage points, NNT = 1/0.04 = 25: treat 25 patients to prevent one event. The RRR of the same data is 67% (4/6), which sounds far larger and is the figure a maker prints; the NNT and ARR are what the patient needs [37][38].
Reading Fig 3's numbers. In Fig 3 the overall diamond sits at RR 7.13 with a 95% CI of 3.92 to 12.95: the CI does not cross 1, so the effect is significant; it is narrow enough (does not span an order of magnitude around 1) to be usefully precise; and the subgroup I² values shown (low, e.g. around 20% and 0%) say the studies within each drug subgroup are consistent enough to pool [40]. A reader who saw only "RR 7.13, p<0.00001" would miss that the CI and I² are what license the conclusion.
Measurement scales are metrics too, judged on three properties [A]: reliability (does it give the same score on repeat/inter-rater, measured by κ or ICC), validity (does it measure what it claims, against an anchor), and responsiveness (does it move when the patient truly changes, tied to the MCID) [12][39]. GAIS is reliable and quick but ceiling-low and evaluator-dependent; validated photonumeric scales are region- and product-specific (a periorbital scale does not grade the nasolabial fold) [13]; FACE-Q is the psychometric benchmark (developed and validated to modern standards) but slower and rarely used by industry [12]. κ (kappa) quantifies agreement beyond chance: below 0.2 poor, 0.4–0.6 moderate, above 0.8 very good; a photonumeric scale with κ under 0.4 between raters is not measuring the face, it is measuring the rater [12].
The metric that belongs in the consultation [A]: the number the patient needs is neither the p value nor the RRR but the absolute one, in "1 in N" form: "about 1 in 300 develop a nodule", "on average 7 in 10 see a visible improvement lasting around X". Natural-frequency phrasing is understood where percentages and relative reductions are not, and it is also the honest framing, because it cannot hide a tiny base rate behind a large relative figure [37][38].
Classic trap: reporting a mean without its dispersion, or an RR without its CI. A point estimate alone hides whether the study knows anything; the spread is the information [36].
B9.7 · Spanish particularity
This block is the evidence and regulatory particularity of Spain; the business, market and autonomic-practice side is covered in B4 (Practice Management) and B7 (Spanish Market), which this section links rather than repeats.
What is specific to Spain and the EU [A]:
| Feature | Spanish/EU particularity | Consequence for appraisal | Ref |
|---|---|---|---|
| No recognised specialty | Aesthetic medicine is not a MIR specialty in Spain; practised by physicians from many backgrounds via master's/CME | Training and its evidence come largely through industry-linked courses | [56] |
| Professional society | SEME issues position and consensus documents, not statutory guidelines | Consensus [A/D]: useful where nothing else exists, inherits panel conflicts |
[56][25] |
| National CPG infrastructure | GuiaSalud (Biblioteca de GPC del SNS) and Osteba develop and appraise CPGs with AGREE II | Almost no aesthetic-specific CPG exists there; the frame is built, the content sparse | [55][25] |
| Medicines regulator | AEMPS, with the CIMA technical-sheet database | CIMA is the medicolegal first source for any medicine (toxin) | [30] |
| Trials of medicines | RD 1090/2015 + EU CTR 536/2014: CEIm approval, CTIS | A toxin study without CEIm/registration is not appraisable | [30][29] |
| Off-label | RD 1015/2009: documented justification, specific consent, non-routine | Most facial toxin/filler use outside the exact IFU indication | [31] |
| Devices | EU MDR 2017/745 applied via AEMPS | Fillers/threads/EBD legal on conformity, not efficacy | [28] |
| Research law | Ley 14/2007 de Investigación Biomédica | Frames human-subject and sample research in Spain | [32] |
| Advertising | Health-advertising rules are largely autonomic (per comunidad autónoma) | Marketing claims are legally constrained but unevenly enforced | [56] |
The aesthetic-evidence gap is not only Spanish, but Spain feels it acutely. The global picture is a field whose published evidence sits low on the ladder: even the plastic-surgery literature that tried to raise its game concedes that most of its studies are small, short and observational, and formal calls "to take evidence-based plastic surgery to the next level" exist precisely because the base level is low [9]. In Spain the gap is widened by structure: with no MIR specialty, the dominant channel for learning a technique is a master's programme or a manufacturer course, so the same body that sells the product frames the evidence a clinician first meets [56]. The corpus of this atlas mirrors it: for a methodology chapter the retrieval returns clinical textbooks, not appraisal standards, because the standards are external and the corpus is a bedside library.
Fig 7. "Is this procedure evidence-based?" – the patient's fair challenge; the professional answer is an [A]–[D] tag plus the honest currency of the source, not "it is proven". (Truswell, 2016, p. 306) [MEDLIB]
> Fuentes: Truswell, Lasers and Light: Peels and Abrasions (2016) [62]
Fig 7 frames the block's practical demand: when a Spanish patient (or a colleague, or an inspector) asks whether a procedure is evidence-based, the defensible answer is the tag and its currency, "this rests on Level 4 consensus, updated 2024, and here is what we do not know", not an appeal to authority [1][9].
How to appraise Spanish and EU sources specifically: 1. CIMA first for any medicine. The ficha técnica is both clinical and medicolegal; it changes and no one alerts you, so re-check quarterly [30]. 2. SEME consensus: read the panel. A consensus of ten speakers from one brand is not a consensus of the field; apply AGREE II domains "stakeholder involvement" and "editorial independence" to it [25][56]. 3. GuiaSalud when a CPG exists. For adjacent domains (scar, wound, laser safety) a GuiaSalud/Osteba CPG appraised with AGREE II outranks any course deck [55][25]. 4. Off-label is the norm, so document it. Under RD 1015/2009 the off-label decision is explicit and consented, which is also your evidence trail [31]. 5. Advertising is regulated per autonomía. A claim legal to publish is not thereby true; separate the regulatory permission from the evidence [56].
The structural consequence of no MIR specialty, in full [A]. Because aesthetic medicine is not a recognised specialty in Spain, there is no MIR training programme, no specialty board setting a curriculum, and no specialty society issuing statutory guidelines; the competency comes through university master's programmes and CME, and the reference documents come through SEME position papers and manufacturer education [56]. Three appraisal consequences follow: (1) the "guideline" a Spanish aesthetic clinician meets is usually a consensus, not an AGREE II-developed CPG, so it must be read as formalised expert opinion carrying its panel's conflicts; (2) the educator is frequently the seller, so the funding-outcome association of B9.5 operates at the level of what techniques are taught, not only what trials are funded [47][56]; (3) the evidence base itself is thin at the source, so external appraisal standards (GRADE, RoB 2, PRISMA) are imported rather than found in-field.
The instruments Spain does have, and how they rank [A]:
- AEMPS + CIMA for any medicine: the ficha técnica is the authoritative, dated, medicolegal source for a toxin's indication, dose and contraindication; it outranks any lecture and it changes without notice, so quarterly re-checking is the rule [30].
- GuiaSalud (Biblioteca de GPC del SNS) and Osteba build and appraise CPGs with AGREE II; for adjacent, evidence-richer domains (wound healing, laser safety, scar management) a GuiaSalud CPG exists and outranks a course deck [55][25].
- SEME consensus and position documents fill the aesthetic-specific space where no CPG exists; useful, but graded as consensus and read with AGREE II domains 2 (stakeholder involvement) and 6 (editorial independence) foremost [56][25].
- CEIm + RD 1090/2015 + EU CTR 536/2014 govern any trial of a medicine done in Spain; a study without CEIm approval and CTIS registration is not appraisable [30][29].
- RD 1015/2009 governs the off-label reality: most facial toxin and filler use sits outside the exact IFU indication, and the decree makes that decision explicit, justified and consented, which doubles as the clinician's evidence trail [31].
- Ley 14/2007 frames biomedical research with human subjects and samples [32].
Worked AGREE II read of a SEME-style consensus (P) [25]: score domain 3 (rigour of development) low if it cites no systematic search; score domain 2 (stakeholder involvement) low if the panel is single-discipline or single-brand; score domain 6 (editorial independence) low if funding and member conflicts are undisclosed. A consensus scoring low on 2, 3 and 6 is not worthless, but it is expert opinion and cannot outrank an AGREE II-developed CPG on the same question, however prestigious its authors [25][56].
Why the gap persists, and it is not only Spanish [A]: the field's evidence sits low on the ladder everywhere, and the discipline's own reformers say so; the call to "take evidence-based plastic surgery to the next level" is a statement that the current level is low, driven by small, short, industry-funded studies [9]. EU MDR 2017/745 is slowly raising the device bar (clinical evaluation and post-market surveillance now demanded), which over time should thicken the filler and device literature, but the back-catalogue was built under a lower bar and most current practice still rests on it [28].
A sustainable evidence routine for a Spanish aesthetic clinic [A] (the cadence that survives a real agenda) [1][30]:
- Weekly, 30 min: PubMed alerts for your 5–10 core topics; read titles, open one or two.
- Monthly: one full guideline or SR on something you do daily.
- Quarterly: re-check CIMA for every medicine you use; technical sheets change and no one alerts you [30].
- Annually: relearn one chapter of your own practice from scratch, as if new.
- At every course: before the first patient, search the technique's evidence; the post-course moment is the peak risk of adopting a fad [9].
The UPO/master's material is a valuable lane with the fastest half-life [MEDLIB]. Master's and course decks ([D]) are where much Spanish aesthetic teaching lives, and they are legitimate for orientation and for stable content (anatomy, mechanism); but a dose or protocol resting on a single teaching slide is never_sufficient_alone, and the deck's currency decays fastest of any lane because it is rarely dated and rarely updated. Treat a slide as a pointer to the primary source, tag it [D], and corroborate before it carries any number [56].
Advertising is regulated but unevenly, and legality is not truth [A]. Health-advertising authorisation in Spain is largely devolved to the comunidades autónomas, so the same claim may face different scrutiny in Barcelona and Madrid; and the EU frames it further through the medicines and MDR advertising rules (a device may not be advertised for an unapproved purpose). The appraisal point is orthogonal to the legal one: a marketing claim that clears the advertising authority has cleared a permission test, not an evidence test, and the [A]–[D] tag is assigned on the evidence, never on the permission [28][56].
The chairside consequence of the gap (P): because the Spanish aesthetic clinician cannot lean on a national CPG for most injectable questions, the burden of appraisal sits with the individual, and the record they keep (Templates 8, 13, 14) is both their evidence trail and their medicolegal defence. The off-label norm (RD 1015/2009) makes this concrete: nearly every facial toxin and filler use is a documented, consented, individually justified decision, which is exactly the discipline B9 exists to install [31][56].
Spain in the EU frame [A]: the medicines route (EU CTR 536/2014, applied through CTIS) and the device route (EU MDR 2017/745) are harmonised across member states, so the evidence bar for a toxin or a filler is, on paper, the same in Barcelona as in Berlin [29][28]. What is not harmonised is the professional structure: some EU countries recognise aesthetic medicine more formally, others (Spain among them) practise it across specialties via master's training, and health-advertising enforcement is national or sub-national. The appraisal takeaway: trust the EU-level product evidence bar as a floor, but do not assume a national "guideline" carries CPG-level rigour, because the body that issued it may be a society, not a guideline developer [55][56].
The Spanish clinician's one-paragraph summary (P): with no MIR specialty and almost no aesthetic-specific CPG, you appraise for yourself; CIMA is your first and medicolegal source for any medicine, SEME consensus is read as expert opinion with AGREE II, GuiaSalud outranks any course deck where a CPG exists, off-label is the documented norm under RD 1015/2009, and the master's-course slide is a pointer, never a proof. The structure that pushes education toward industry is exactly why the appraisal discipline of B9.1 to B9.6 is not optional in Spain.
Pharmacovigilance and device-vigilance are an under-used Spanish evidence source [A]. Adverse events with a medicine (toxin) are reportable to the AEMPS through the pharmacovigilance system, and device incidents (filler granuloma, device burn) through device-vigilance; these datasets are exactly the long-horizon, rare-event signal the RCT literature lacks, and MDR post-market surveillance is strengthening the device side [28][30]. For a Spanish clinician the practical loop is twofold: report your own serious events (a legal duty and a data contribution), and read the aggregated safety communications as a real, if coarse, evidence lane on late complications that no sponsored trial will surface [28][53].
Classic trap: treating a SEME or master's-course consensus as if it were a GuiaSalud CPG. They are different instruments: one is expert opinion formalised (inherits its panel's conflicts), the other is a methodologically appraised guideline; grade them accordingly and never let the course deck outrank the appraised CPG [25][55].
B9.8 · Organizational alternatives
When the RCT ladder cannot answer, the answer is not to abandon rigour but to change how evidence is organised. The alternatives, and what each repairs [A]:
| Alternative | What it repairs | Aesthetic use case | Ref |
|---|---|---|---|
| Real-world evidence / registries | No RCT on long-term or rare outcomes | Filler-registry late nodules, device adverse-event tracking | [53] |
| Living systematic review / living guideline | Evidence dating between updates | Fast-moving areas (biostimulators, exosomes) | [52] |
| N-of-1 trial | Individual, not population, answer | Does this patient respond to this regimen | [54] |
| Pragmatic trial | RCT external validity too narrow | Effectiveness in ordinary clinics, ordinary injectors | [9] |
| Core outcome set (COS) | Endpoint chaos blocks pooling | Agree one validated outcome (e.g. FACE-Q) across trials | [12] |
| GRADE-ADOLOPMENT | No local guideline, no time for one | Adopt/adapt an existing appraised CPG to your setting | [8] |
| New evidence pyramid | The layer model is too rigid | Read SR as the lens, quality moves a study up/down | [4] |
Real-world evidence and registries. RWE is data from routine care (registries, electronic records, claims, wearables) analysed to answer questions an RCT never will: what happens at 5 or 15 years, how a device performs across thousands of ordinary users, which rare adverse events cluster [53]. For aesthetics, where the orphan questions are exactly the long-horizon and rare-event ones, a well-run product registry is often the strongest feasible evidence, and EU MDR post-market surveillance is pushing device makers toward it [28][53]. The caveat: RWE without a comparator and without confounding control is a large case series, so appraise it with ROBINS-I, not with the awe its sample size invites [22].
Living systematic reviews and living guidelines. A living SR is continuously updated as new studies appear, rather than frozen at publication and stale within two years [52]. It matters most where a field moves fast (regenerative and biostimulator aesthetics), and it changes the reader's job: check the "last updated" date the way you check a technical sheet.
N-of-1 trials. A single patient is cycled, ideally blinded and randomised, through the candidate treatment and its control across multiple periods; the patient is their own control, and the answer is about them specifically [54]. In aesthetics it fits chronic, self-assessed, reversible interventions (a topical, a device cadence) better than an irreversible injection.
Core outcome sets and pragmatic trials. Much aesthetic evidence cannot be pooled because every trial invents its own endpoint; a COS fixes a minimum validated outcome (FACE-Q and its modules are the candidate) so trials become comparable [12]. Pragmatic trials relax the artificial conditions of explanatory RCTs (expert injectors, ideal patients) to measure effectiveness in the ordinary clinic, trading internal for external validity, which is the validity aesthetics most lacks [9].
GRADE-ADOLOPMENT and the wavy pyramid. Rather than build a guideline from zero, ADOPT an existing appraised CPG, or ADAPT/develop de novo only the recommendations that differ locally, keeping the GRADE audit trail [8]. And the mental model to retire is the rigid pyramid: Murad's revision makes the inter-layer lines wavy (study quality moves a study between layers) and lifts the systematic review out of the stack to be the lens through which every layer is viewed [4].
The four schools of EBM, and where each fails [A] (the position that governs this atlas is Sackett's):
| Stance | Thesis | Where it fails |
|---|---|---|
| MBE fundamentalist | No RCT, no action | Paralyses ~90 % of practice, hyaluronidase in occlusion included; confuses absence of evidence with evidence of absence [1] |
| Experientialist | "20 years doing it, it works" | Experience does not control confirmation bias or see who never returns; the dissatisfied do not come back, so the personal series is biased by construction [11] |
| Pragmatic EBM ✅ | Best available evidence + expertise + patient values, grade stated | Sackett's original definition, and the one this atlas adopts [1] |
| Consensualist | Follow expert consensus | Useful where nothing else exists; inherits the panel's conflicts whole [47] |
The operational synthesis (P): acting on low evidence is legitimate and often obligatory in aesthetics; acting without knowing the evidence is low is not. The alternatives above exist so that "there is no RCT" stops being the end of the sentence and becomes the start of choosing the right instrument.
More organisational forms, and when each earns its place [A]:
| Form | What it is | When it beats a single RCT |
|---|---|---|
| Umbrella review | A review of systematic reviews on a topic | When many SRs exist and disagree; ranks them by AMSTAR-2 [23] |
| Scoping review | Maps the extent and gaps of a literature | When the question is "what has been studied at all?" |
| Prospective meta-analysis | Trials pre-agree to pool before results | Removes the publication-bias distortion of retrospective pooling [41] |
| Adaptive / platform trial | Design changes by pre-planned rules; arms added/dropped | When many products compete and a fixed RCT is too slow |
| Target-trial emulation | Analyse observational data as the RCT you cannot run | When RWE is all there is; forces explicit confounding control [22] |
| Bayesian analysis | Updates a prior with the data to a posterior | When priors are informative and n is small (much of aesthetics) |
| PRO registry | Longitudinal patient-reported outcomes (FACE-Q) at scale | Captures the patient's own long-horizon verdict [12] |
Target-trial emulation is the discipline that makes real-world evidence defensible: before touching the data, specify the RCT you would have run (eligibility, treatment strategies, assignment, outcome, follow-up, analysis), then emulate each element in the observational data, which surfaces the confounding and immortal-time traps that a naive registry analysis hides [22][53]. Bayesian methods fit aesthetics structurally: with small samples and informative priors (mechanism, prior products in the same class), a posterior probability ("85% chance this filler outlasts that one") is often more honest than a dichotomous p value from n=30, provided the prior is stated and defensible [35].
Prospective and platform designs fix the field's two worst distortions [A]: publication bias (a prospective meta-analysis binds trials to publish and pool regardless of result [41]) and glacial comparison (an adaptive platform lets many products share a control arm and drop losers by rule, instead of a decade of underpowered head-to-heads [9]). Neither is common in aesthetics yet; naming them is naming the direction the field's evidence must move.
The schools, with the medicolegal and practical reading [A]: the fundamentalist who refuses to act without an RCT would withhold hyaluronidase during an occlusion, which is indefensible clinically and legally [1]. The experientialist who trusts twenty years of memory is exposed the day a late complication he never tracked surfaces in a claim, because he kept no denominator (B9.4, Template 14) [11]. The consensualist who follows a single-brand consensus inherits its conflicts and cannot show independent judgement if challenged [47]. Sackett's pragmatic EBM, best available evidence integrated with expertise and the patient's values, with the grade stated, is the only stance that is simultaneously good practice and a defensible record [1].
Choosing the alternative, a decision path [A]: is the question about a rare or long-horizon outcome? → registry/RWE with target-trial emulation [22][53]. Is the field moving fast? → living SR/guideline [52]. Is the question about this individual, and the treatment reversible? → N-of-1 [54]. Is the problem that no two trials measured the same thing? → adopt a core outcome set [12]. Is there no local guideline and no time? → GRADE-ADOLOPMENT an existing one [8]. Do many products need comparing at once? → an adaptive platform trial [9]. The alternative is chosen by the shape of the gap, not by preference.
Grading the certainty of real-world and prognostic evidence [A]. GRADE was extended so that observational and prognostic bodies are not dismissed as "low" by default: RWE can be upgraded for a large effect, a dose-response gradient, or when all plausible confounding would shrink the observed effect, exactly the three signals that rescued hyaluronidase in B9.1 [6][53]. So a well-run filler registry showing a large, dose-related late-nodule signal can reach moderate certainty, which a naive "it is only observational" dismissal would miss. Appraise the registry with ROBINS-I, then grade the body with GRADE; the two steps are different [22][6].
Know the review types apart [A]: a systematic review answers a focused question with a reproducible search and risk-of-bias appraisal [15]; a scoping review maps what exists and where the gaps are; an umbrella review reviews the systematic reviews and ranks them by AMSTAR-2 [23]; a rapid review trades some method for speed and must declare what it cut. Calling a hand-picked narrative a "review" is the commonest label inflation; the PRISMA flow (Fig 1) is what separates a systematic review from a reading list [15][16].
A worked living-guideline read (P): for a fast-moving topic (biostimulator injectables), a living guideline updated within the last months and carrying a GRADE certainty per recommendation outranks a five-year-old static consensus, even a prestigious one, because currency is part of validity for a changing claim (B9.1); the reader's first move is to find the "last updated" date, exactly as with a technical sheet [52][30].
How aesthetics should reorganise its evidence (P): the field's specific failures (short follow-up, endpoint chaos, publication bias, industry funding, tiny samples) map onto specific fixes (registries for the long horizon, core outcome sets for the endpoints, prospective meta-analysis for the bias, pragmatic and platform trials for effectiveness, Bayesian analysis for the small samples). None requires abandoning rigour; each is rigour reorganised for a field that cannot run the classic RCT ladder. Naming them turns "there is no RCT" from an excuse into a menu.
Core outcome sets are the fix for the endpoint chaos [A]. The reason aesthetic trials cannot be pooled is that each invents its own outcome (one uses GAIS at 12 weeks, the next patient satisfaction at 6 months, the next 3D volumetry at a year), so a meta-analysis mixes measures that do not mean the same thing and returns a high-I² number describing nothing [40]. A core outcome set is a minimum agreed list of what every trial in a domain must measure, ideally a validated PRO such as FACE-Q plus one objective measure and a defined harms set; once trials share it, pooling becomes meaningful and the field can finally compare products head to head [12]. Adopting a COS is the cheapest structural upgrade available to aesthetic evidence, and it requires no new trial, only agreement.
The evidence-to-decision bridge still applies to every alternative [A]: whatever the evidence's organisational form (RCT, registry, N-of-1, living guideline), the move to a recommendation runs through the same GRADE Evidence-to-Decision logic (benefit vs harm, certainty, patient values, resources) [7][8]. The alternatives change how certainty is earned, never the fact that a recommendation must weigh values and harms explicitly; a registry signal of a rare severe harm can drive a strong recommendation against a product on low certainty, exactly mirroring the hyaluronidase logic of B9.1 in the opposite direction.
Classic trap: treating a large registry or an RWE dataset as if its size were rigour. A million records with no comparator and unadjusted confounding is a case series at scale; appraise it as one [22][53].
Coverage vs UPO
The UPO máster teaches research methodology chiefly through the TFM (trabajo fin de máster) elaboration criteria and scattered statistics in the antienvejecimiento module; it has no dedicated critical-appraisal unit. What it teaches, its state here, and what the atlas adds:
| UPO topic taught | State in this chapter | What the atlas adds |
|---|---|---|
| TFM elaboration criteria (how to structure a thesis) | Covered in B9.3 (PICO, search) and B9.4 (templates) | The full 5A cycle and design-matched appraisal tools UPO omits |
| Basic descriptive statistics (antiaging module) | Covered in B9.6 (mean/SD, CI, Fig 6) | RR/OR/HR, ARR/NNT, LR, I², MCID with reference values |
| Industry/COI as a lecture aside (sugar-industry slide, Fig 2) | Covered and quantified in B9.5 | Lundh's RR 1.27/1.34 sponsorship effect; the six mechanisms |
| "Levels of evidence" mentioned generically | Covered in B9.1 | OCEBM 2011 + GRADE, and why never to blend them |
Rows UPO does not cover at all (each is an unclosed gap the atlas fills):
| Not in UPO | Where the atlas covers it |
|---|---|
| GRADE certainty vs strength of recommendation | B9.1 (worked on hyaluronidase) |
| RoB 2 / ROBINS-I / AMSTAR-2 / AGREE II instruments | B9.3, B9.4 (templates 3, 11, 12) |
| EQUATOR reporting guidelines (CONSORT/PRISMA/STROBE/STARD) | B9.2 (with Fig 1 PRISMA flow) |
| Reading a forest plot / funnel plot / I² | B9.6 (Figs 3, 4) |
| EU MDR vs medicines route, and the device evidence bar | B9.2, B9.7 |
| Spanish regulatory frame (RD 1090/2015, RD 1015/2009, AEMPS) | B9.7 |
| Organisational alternatives (RWE, living SR, N-of-1, COS) | B9.8 |
| Diagnostic metrics (sens/spec, PPV/NPV, LR, pre/post-test) | B9.6 (worked 2×2) |
UPO material ([D]) is the fastest-ageing lane: it is a pointer to the primary source, never a proof, and a number resting on a single slide is never_sufficient_alone (B9.7). No UPO methodology topic surfaced by the retrieval is left uncovered here.
Self-assessment
Ten active-recall questions built only from facts already in this chapter. Answers folded.
- What two things does GRADE hold apart that the classical hierarchy blurs, and what are their levels?
Answer
**Certainty of evidence** (high / moderate / low / very low) and **strength of recommendation** (strong / conditional). A strong recommendation on low certainty is legitimate (hyaluronidase in occlusion). [B9.1]- Name the five GRADE downgrade domains and the three upgrade domains.
Answer
Downgrade: risk of bias, inconsistency, indirectness, imprecision, publication bias. Upgrade (observational): large effect, dose-response gradient, all plausible confounding would reduce the observed effect. [B9.1]- Industry-sponsored studies reach favorable efficacy results and conclusions at what risk ratios, and does a clean risk-of-bias assessment clear that bias?
Answer
Favorable efficacy RR 1.27 (95% CI 1.17 to 1.37); favorable conclusions RR 1.34 (1.19 to 1.51). No: the sponsorship effect survives a clean RoB 2, so conflicts are judged separately, first. [B9.5]- On a forest plot, what do the square area, the horizontal line, the vertical line and the diamond each represent?
Answer
Square area = the study's weight; horizontal line = that study's 95% CI (crossing the null = non-significant); vertical line = no effect (RR/OR = 1); diamond = pooled estimate, its width the pooled CI. [B9.6, Figs 3–4]- What are the Cochrane I² bands, and above what value is pooling questionable?
Answer
0–40% may be unimportant, 30–60% moderate, 50–90% substantial, 75–100% considerable; above ~50% pooling is questionable. [B9.6]- A step drops a complication from 6% to 2%. Give the ARR, the NNT and the RRR.
Answer
ARR = 4 percentage points; NNT = 1/0.04 = 25; RRR = 4/6 = 67% (the figure a maker prints). [B9.6]- In 1000 patients, 10% prevalence, a 90%/90% test: what is the PPV, and why?
Answer
TP=90, FP=90, so PPV = 90/180 = 50%: at low prevalence most positives are false, which is why PPV cannot be quoted without prevalence and the LR is the portable number. [B9.6]- Which product classes are medicines and which are devices, and how does that change the evidence bar?
Answer
Botulinum toxin is a medicine (AEMPS/EMA route, registration RCTs against an authorised indication); fillers, threads and most EBD are devices under MDR 2017/745 (CE mark on conformity, historically no efficacy RCT required). [B9.2, B9.7]- Name the five RoB 2 domains, and which one governs the overall judgement.
Answer
D1 randomisation, D2 deviations from intended intervention, D3 missing outcome data, D4 outcome measurement, D5 selective reporting; the worst single domain governs the overall verdict. [B9.3, B9.4]- Why is a split-face trial's data wrong if the two hemifaces are counted as two independent patients, and what off-label decree governs most facial toxin/filler use in Spain?
Answer
Two hemifaces of one patient are not two patients; counting them independently doubles the apparent n and inflates significance, so a paired analysis is required. Off-label use is governed by RD 1015/2009 (documented justification, specific consent, non-routine). [B9.5, B9.7]What's new and trends (2023–2026)
| Period | What changed | Maturity | Atlas ref |
|---|---|---|---|
| 2019→ | RoB 2 and ROBINS-I became the Cochrane standard, replacing the older RoB 1 tool | clinically actionable now | B9.3, B9.4 [21][22][24] |
| 2021→ | PRISMA 2020 (new flow diagram) superseded PRISMA 2009 as the SR reporting norm | clinically actionable now | B9.2 [15][16] |
| 2021→ | EU MDR 2017/745 in full application: clinical evaluation and post-market surveillance demanded of devices | clinically actionable now | B9.2, B9.7 [28] |
| 2023→ | Living systematic reviews and living guidelines moved to accepted practice | clinically actionable now | B9.8 [52] |
| 2023→ | Real-world evidence and target-trial emulation entering mainstream regulatory use | promising but not validated | B9.8 [22][53] |
| 2023→ | FACE-Q consolidated as the psychometric PRO benchmark for facial aesthetics | clinically actionable now | B9.2, B9.6 [12] |
| 2016→ | GRADE Evidence-to-Decision frameworks became the standard bridge to a recommendation | clinically actionable now | B9.1, B9.4 [7][8] |
| 2024→ | LLM-assisted screening and risk-of-bias tools appearing in appraisal workflows | preclinical/speculative | B9.3 [21] |
| ongoing | "Clinically proven" and "X-month result" device marketing outrunning its evidence | unsupported commercial claim | B9.5 [47] |
What did NOT change, and why the old references still stand. The core of critical appraisal is stable, so several load-bearing references here are deliberately old and remain state of the art: the logic of the p value and the confidence interval (ASA statement, 2016; Gardner-Altman, 1986) is unchanged [35][36]; the number needed to treat (Cook-Sackett, 1995) is defined as it was [37]; I² (Higgins-Thompson, 2003) still measures heterogeneity the same way [40]; Sackett's 1996 definition of evidence-based medicine is still the one this atlas adopts [1]; and the funding-outcome association (Lundh, 2017) has been stable across successive Cochrane updates, not overturned [47]. The field's structural problem is also unchanged: the aesthetic evidence base remains low on the ladder, exactly as described a decade ago [9]. A stable method does not need a 2025 citation to be current; the year is visible, and the claim is of the stable class (B9.1).
Unexplored directions (AI speculation)
> Speculative section. The items below are model-generated research directions, not evidence and not recommendations. Each is tagged [IA-ESPEC], states the cited anchor it grows from, the proposal, and what would settle it. None contains a dose, a product or a protocol a reader could act on.
-
[IA-ESPEC]Independent-funding pool for the orphan questions. Anchor: sponsored trials reach favorable conclusions at RR 1.34, and the field's long-horizon questions (15-year outcomes, cumulative load, "does treating less do better") have no literature because no maker funds them [47] (B9.5). Proposal: a levy on device/toxin sales funding an independent body to run exactly those trials. Expected effect: the RR 1.34 conclusion-favorability gap narrows toward 1.0 under independent funding. Confounder: differences in trial quality and product mix between funders (match on design and indication). What would settle it: compare effect estimates and conclusion-direction on the same products between levy-funded and industry-funded trials; convergence would refute the need, divergence would quantify it. -
[IA-ESPEC]A mandated aesthetic core outcome set. Anchor: trials cannot be pooled because each invents its own endpoint, so meta-analyses return uninterpretable high-I² numbers [40] (B9.8). Proposal: journals require a defined aesthetic COS (a validated PRO plus one objective measure plus a harms set) as a submission condition. Expected effect: a measurable fall in outcome-attributable heterogeneity after adoption. Confounder: secular improvement in trial quality unrelated to the COS (interrupted time-series with a control set). What would settle it: measure the share of meta-analytic heterogeneity attributable to outcome differences before and after adoption; a fall would confirm the mechanism. -
[IA-ESPEC]Shared EU registry analysed by target-trial emulation. Anchor: real-world evidence is the strongest feasible design for rare and long-horizon aesthetic outcomes, but naive registry analysis hides confounding and immortal-time bias [22][53] (B9.8). Proposal: a common EU filler/device registry whose default analysis is a pre-specified target-trial emulation. Expected effect: emulated estimates fall within the CI of the reference RCTs. Confounder: residual confounding by indication and injector skill (encoded in the target-trial eligibility). What would settle it: calibrate emulated estimates against the few existing RCTs on the same products; agreement would validate the approach, disagreement would locate the residual confounding. -
[IA-ESPEC]Automated currency-decay flagging of claims. Anchor: claims split into stable and changing, and changing claims (regulation, IFU, device parameters) decay with time while their tag does not [1] (B9.1). Proposal: an automated flag that downgrades a changing claim's visibility as elapsed time since its source grows. Expected effect: flagged-stale claims show a higher later contradiction rate than unflagged. Confounder: topic churn (some topics change faster regardless; stratify by claim class). What would settle it: test whether flagged-stale claims have a higher later contradiction or revision rate than unflagged ones; a positive association would justify the flag. -
[IA-ESPEC]Aggregated N-of-1 networks for reversible interventions. Anchor: the N-of-1 design answers the individual question and suits chronic, reversible, self-rated aesthetic interventions [54] (B9.8). Proposal: pool many single-patient trials into a network to build a population estimate without a classic parallel RCT. Expected effect: the pooled N-of-1 estimate coincides with the parallel RCT estimate. Confounder: period carryover (enforce washout; test order effects). What would settle it: run a parallel RCT and an aggregated N-of-1 series on the same reversible intervention and test whether their estimates coincide. -
[IA-ESPEC]LLM-assisted risk-of-bias screening as a first pass. Anchor: RoB 2 makes a judgement reproducible by naming the failing domain, but applying it by hand across a literature is slow [21] (B9.3). Proposal: a language model performs a first-pass RoB 2 domain screen that a human then adjudicates. Expected effect: model-expert agreement (κ) exceeds a pre-set floor. Confounder: the model echoing the paper's own spin (blind it to the abstract's conclusions). What would settle it: measure inter-rater agreement (κ) between model and expert across a labelled corpus and test whether disagreements cluster in one domain (e.g. D5 selective reporting), which would bound where the tool is trustworthy.
§ Safety
Critical appraisal has its own safety rules: the errors of B9 do not spill blood directly, but they put the wrong dose, the missed late complication and the uninformed consent into the room. The non-negotiable rules:
| Safety rule | Why | If violated |
|---|---|---|
| Never a dose, threshold, PMID or DOI from memory | Recalled identifiers and figures are a documented failure mode | A fabricated dose or a citation that does not resolve [1] |
(P) and [MODELO] never carry a number |
The model's reasoning is not evidence | A quantitative claim with no traceable source |
A dose on a single UPO slide is never_sufficient_alone |
The teaching lane ages fastest and is undated | A stale or wrong parameter acted on [56] |
| Read the follow-up before "well tolerated" | 12-week studies cannot see 6–24 month nodules | A late complication the patient was told could not happen [10] |
| Off-label use documented under RD 1015/2009 | Explicit justification + specific consent | Medicolegal exposure and an uninformed patient [31] |
| CIMA/IFU is the medicolegal first source, re-checked quarterly | Technical sheets change without notice | Acting on a superseded dose or contraindication [30] |
| Give the absolute risk in consent, in natural frequencies | Relative figures hide the base rate | Consent that is legally present but not informed [37][38] |
| Keep conflicting values whole, never averaged | Averaging a 1:2.5 and a 1:3 invents a third number | A dose no source supports |
| Report serious events to pharmacovigilance/device-vigilance | Rare/late signals live only in aggregate | The field never learns of the harm [28] |
| Distrust a switched primary outcome | Outcome switching manufactures significance | A benefit believed that the trial did not show [45] |
The one rule under all of them (P): acting on low evidence is legitimate in aesthetics and often obligatory; acting without knowing the evidence is low, and without telling the patient so, is not. The tag is the safety device: a claim without a grade is not finished, and a chapter that contradicts the tag rules of B9.1 has the error in the chapter, not in B9 [1][9].
Anti-leakage note. Every quantitative claim in this chapter traces to a retrieved corpus source ([MEDLIB], opened before captioning) or an external reference with a verified DOI ([A]/[B]). Where the corpus does not cover a methodology facet (OCEBM, GRADE, the EQUATOR instruments, EU/Spanish regulation), the content is external-lane and cited as such; no facet was filled from model memory. The retrieval that grounds this is evaluation/runs/B9.1–B9.5.jsonl; the methodology authorities are external because a bedside corpus does not contain them, which is stated, not hidden.
References
Evidence tier in brackets: [A] guideline/consensus/SR · [B] primary study · [C] monograph · [D] slide/course · [MEDLIB] own corpus. External identifiers are verified and linked.
- Sackett DL, Rosenberg WMC, Gray JAM, Haynes RB, Richardson WS. Evidence based medicine: what it is and what it isn't. BMJ. 1996;312:71-72.
[A]DOI 10.1136/bmj.312.7023.71 - OCEBM Levels of Evidence Working Group (Howick J, et al.). The Oxford 2011 Levels of Evidence. Centre for Evidence-Based Medicine, University of Oxford; 2011.
[A]cebm.ox.ac.uk - Guyatt GH, Oxman AD, Vist GE, Kunz R, Falck-Ytter Y, Alonso-Coello P, Schünemann HJ. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924-926.
[A]DOI 10.1136/bmj.39489.470347.AD - Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evid Based Med. 2016;21:125-127.
[A]DOI 10.1136/ebmed-2016-110401 - Guyatt GH, Oxman AD, Kunz R, et al. GRADE guidelines: 1. Introduction. J Clin Epidemiol. 2011;64:383-394.
[A]DOI 10.1016/j.jclinepi.2010.04.026 - Balshem H, Helfand M, Schünemann HJ, et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol. 2011;64:401-406.
[A]DOI 10.1016/j.jclinepi.2010.07.015 - Andrews J, Guyatt G, Oxman AD, et al. GRADE guidelines: 15. Going from evidence to recommendation. J Clin Epidemiol. 2013;66:726-735.
[A]DOI 10.1016/j.jclinepi.2013.02.003 - Alonso-Coello P, Schünemann HJ, Moberg J, et al. GRADE Evidence to Decision (EtD) frameworks. BMJ. 2016;353:i2016.
[A]DOI 10.1136/bmj.i2016 - Eaves FF 3rd, Rohrich RJ, Sykes JM. Taking evidence-based plastic surgery to the next level. Aesthet Surg J. 2013;33:459-465.
[B]DOI 10.1177/1090820X13493766 - Alam M (ed). Evidence-Based Procedural Dermatology. Springer; 2019.
[C][MEDLIB] - Ioannidis JPA. Why most published research findings are false. PLoS Med. 2005;2:e124.
[B]DOI 10.1371/journal.pmed.0020124 - Klassen AF, Cano SJ, Schwitzer JA, Scott AM, Pusic AL. Development and psychometric evaluation of the FACE-Q scales. JAMA Facial Plast Surg. 2016;18:113-121.
[B]DOI 10.1001/jamafacial.2015.1445 - Kim P, Ahn JT. A validated rating scale for hyperkinetic facial lines. Arch Facial Plast Surg. 2004;6:253-256.
[B]DOI 10.1001/archfaci.6.4.253 - Schulz KF, Altman DG, Moher D; CONSORT Group. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c332.
[A]DOI 10.1136/bmj.c332 - Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.
[A]DOI 10.1136/bmj.n71 - Moher D, Liberati A, Tetzlaff J, Altman DG; PRISMA Group. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. BMJ. 2009;339:b2535.
[A]DOI 10.1136/bmj.b2535 - von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, Vandenbroucke JP; STROBE Initiative. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement. Lancet. 2007;370:1453-1457.
[A]DOI 10.1016/S0140-6736(07)61602-X - Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527.
[A]DOI 10.1136/bmj.h5527 - Chan AW, Tetzlaff JM, Altman DG, et al. SPIRIT 2013 statement: defining standard protocol items for clinical trials. Ann Intern Med. 2013;158:200-207.
[A]DOI 10.7326/0003-4819-158-3-201302050-00583 - Ioannidis JPA, Evans SJW, Gøtzsche PC, et al. Better reporting of harms in randomized trials: an extension of the CONSORT statement. Ann Intern Med. 2004;141:781-788.
[A]DOI 10.7326/0003-4819-141-10-200411160-00009 - Sterne JAC, Savović J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
[A]DOI 10.1136/bmj.l4898 - Sterne JA, Hernán MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
[A]DOI 10.1136/bmj.i4919 - Shea BJ, Reeves BC, Wells G, et al. AMSTAR 2: a critical appraisal tool for systematic reviews. BMJ. 2017;358:j4008.
[A]DOI 10.1136/bmj.j4008 - Higgins JPT, Altman DG, Gøtzsche PC, et al. The Cochrane Collaboration's tool for assessing risk of bias in randomised trials. BMJ. 2011;343:d5928.
[A]DOI 10.1136/bmj.d5928 - Brouwers MC, Kho ME, Browman GP, et al. AGREE II: advancing guideline development, reporting and evaluation in health care. CMAJ. 2010;182:E839-E842.
[A]DOI 10.1503/cmaj.090449 - World Medical Association. WMA Declaration of Helsinki: ethical principles for medical research involving human subjects. JAMA. 2013;310:2191-2194.
[A]DOI 10.1001/jama.2013.281053 - De Angelis C, Drazen JM, Frizelle FA, et al. Clinical trial registration: a statement from the ICMJE. N Engl J Med. 2004;351:1250-1251.
[A]DOI 10.1056/NEJMe048225 - Regulation (EU) 2017/745 of the European Parliament and of the Council on medical devices (MDR). Off J Eur Union. 2017.
[A]EUR-Lex 32017R0745 - Regulation (EU) No 536/2014 on clinical trials on medicinal products for human use. Off J Eur Union. 2014.
[A]EUR-Lex 32014R0536 - Real Decreto 1090/2015, de 4 de diciembre, por el que se regulan los ensayos clínicos con medicamentos. BOE núm. 307; 2015.
[A]BOE-A-2015-14082 - Real Decreto 1015/2009, de 19 de junio, por el que se regula la disponibilidad de medicamentos en situaciones especiales (uso off-label). BOE núm. 174; 2009.
[A]BOE-A-2009-12002 - Ley 14/2007, de 3 de julio, de Investigación Biomédica. BOE núm. 159; 2007.
[A]BOE-A-2007-12945 - EQUATOR Network. Enhancing the QUAlity and Transparency Of health Research: library of reporting guidelines.
[A]equator-network.org - Critical Appraisal Skills Programme (CASP). CASP checklists (RCT, systematic review, cohort, diagnostic).
[A]casp-uk.net - Wasserstein RL, Lazar NA. The ASA statement on p-values: context, process, and purpose. Am Stat. 2016;70:129-133.
[A]DOI 10.1080/00031305.2016.1154108 - Gardner MJ, Altman DG. Confidence intervals rather than P values: estimation rather than hypothesis testing. BMJ. 1986;292:746-750.
[A]DOI 10.1136/bmj.292.6522.746 - Cook RJ, Sackett DL. The number needed to treat: a clinically useful measure of treatment effect. BMJ. 1995;310:452-454.
[A]DOI 10.1136/bmj.310.6977.452 - Naylor CD, Chen E, Strauss B. Measured enthusiasm: does the method of reporting trial results alter perceptions of therapeutic effectiveness? Ann Intern Med. 1992;117:916-921.
[B]DOI 10.7326/0003-4819-117-11-916 - Jaeschke R, Singer J, Guyatt GH. Measurement of health status: ascertaining the minimal clinically important difference. Control Clin Trials. 1989;10:407-415.
[B]DOI 10.1016/0197-2456(89)90005-6 - Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327:557-560.
[B]DOI 10.1136/bmj.327.7414.557 - Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315:629-634.
[B]DOI 10.1136/bmj.315.7109.629 - Altman DG, Bland JM. Diagnostic tests 1: sensitivity and specificity. BMJ. 1994;308:1552.
[A]DOI 10.1136/bmj.308.6943.1552 - Deeks JJ, Altman DG. Diagnostic tests 4: likelihood ratios. BMJ. 2004;329:168-169.
[A]DOI 10.1136/bmj.329.7458.168 - Hollis S, Campbell F. What is meant by intention to treat analysis? Survey of published randomised controlled trials. BMJ. 1999;319:670-674.
[B]DOI 10.1136/bmj.319.7211.670 - Chan AW, Hróbjartsson A, Haahr MT, Gøtzsche PC, Altman DG. Empirical evidence for selective reporting of outcomes in randomized trials. JAMA. 2004;291:2457-2465.
[B]DOI 10.1001/jama.291.20.2457 - Boutron I, Dutton S, Ravaud P, Altman DG. Reporting and interpretation of randomized controlled trials with statistically nonsignificant results for primary outcomes. JAMA. 2010;303:2058-2064.
[B]DOI 10.1001/jama.2010.651 - Lundh A, Lexchin J, Mintzes B, Schroll JB, Bero L. Industry sponsorship and research outcome. Cochrane Database Syst Rev. 2017;2:MR000033.
[A]DOI 10.1002/14651858.MR000033.pub3 - Kearns CE, Schmidt LA, Glantz SA. Sugar industry and coronary heart disease research: a historical analysis of internal industry documents. JAMA Intern Med. 2016;176:1680-1685.
[B]DOI 10.1001/jamainternmed.2016.5394 - Simmons JP, Nelson LD, Simonsohn U. False-positive psychology: undisclosed flexibility in data collection and analysis. Psychol Sci. 2011;22:1359-1366.
[B]DOI 10.1177/0956797611417632 - Barnett AG, van der Pols JC, Dobson AJ. Regression to the mean: what it is and how to deal with it. Int J Epidemiol. 2005;34:215-220.
[B]DOI 10.1093/ije/dyh299 - Lesaffre E, Philstrom B, Needleman I, Worthington H. The design and analysis of split-mouth studies: what statisticians and clinicians should know. Stat Med. 2009;28:3470-3482.
[B]DOI 10.1002/sim.3634 - Elliott JH, Synnot A, Turner T, et al. Living systematic review: 1. Introduction. J Clin Epidemiol. 2017;91:23-30.
[A]DOI 10.1016/j.jclinepi.2017.08.010 - Sherman RE, Anderson SA, Dal Pan GJ, et al. Real-world evidence: what is it and what can it tell us? N Engl J Med. 2016;375:2293-2297.
[A]DOI 10.1056/NEJMsb1609216 - Guyatt G, Sackett D, Taylor DW, Chong J, Roberts R, Pugsley S. Determining optimal therapy: randomized trials in individual patients (N-of-1). N Engl J Med. 1986;314:889-892.
[B]DOI 10.1056/NEJM198604033141406 - GuiaSalud. Biblioteca de Guías de Práctica Clínica del Sistema Nacional de Salud (España).
[A]guiasalud.es - Sociedad Española de Medicina Estética (SEME). Documentos de consenso y posicionamiento en medicina estética. Madrid: SEME; 2024.
[A]seme.org - Manchikanti L, et al. Essentials of Interventional Techniques in Managing Chronic Pain. Springer; 2024. ISBN 9783031462160.
[C][MEDLIB](PRISMA 2020 flow, Fig 1) - Barash PG, et al. Clinical Anesthesia. 5th ed; 2006.
[C][MEDLIB](forest plot, Fig 3) - Cross ME, Plunkett EVE. Physics, Pharmacology and Physiology for Anaesthetists. Cambridge; 2008. ISBN 9780521700443.
[C][MEDLIB](forest-plot schematic, Fig 4) - Schmidt RF, Willis WD (eds). Encyclopedia of Pain. Springer; 2007. ISBN 9783540439578.
[C][MEDLIB](effect-size figure, Fig 5) - Aitkenhead AR, Rowbotham DJ, Smith G. Textbook of Anaesthesia. 4th ed; 2001. ISBN 0443063818.
[C][MEDLIB](descriptive statistics, Fig 6) - Truswell WH IV. Lasers and Light: Peels and Abrasions. Thieme; 2016.
[C][MEDLIB](evidence-based cartoon, Fig 7)
Verification: B9 · 2026-08-24 · author atlas-chapter (fresh context). Scope: the PRACTICA 8-block frame mapped onto evidence-based practice and critical appraisal; the governing chapter for the atlas [A–D] tag system. Corpus pass: evaluation/runs/B9.1–B9.5.jsonl (5 subchapters, all facets, k=8, figure-k=6, generic+aesthetic-regenerative overlay; VERDICT usable). B9 is a methodology chapter: the corpus returned clinical textbooks (top hits Manchikanti, Gullo, Fishman), so the methodology authorities (OCEBM, GRADE, RoB 2/ROBINS-I/AMSTAR-2, EQUATOR, EU MDR/CTR, Spanish RD, SEME) are external lane, every DOI verified via Crossref and the Lundh sponsorship effect-sizes via Europe PMC; the corpus served as the figure source and for [C] corroboration. Figures: 7, each opened with Read before captioning (PRISMA flow, real + schematic forest plots, effect-size distributions, descriptive-statistics page, sugar-industry COI paper, evidence-based cartoon); figure_pick receipt at _images/B9/figure-pick-receipt.json. Salvage: prior ES version (5 subchapters) salvaged whole; no language-neutral fact dropped. Anti-leakage: no facet filled from memory; (P)/[MODELO] carry no dose or quantitative claim. Added beyond the 5 old subchapters: the chapter expands the retired 5-subchapter ES version into the current 8-block PRACTICA frame, adding the regulatory/normative block (B9.2), the templates block (B9.4), the metrics-with-reference-values block (B9.6), the Spanish-particularity block (B9.7) and the organizational-alternatives block (B9.8), because the restructure requires them and the old file covered none as standalone units.
Navigate
Prev: B8 — Private Practice & Business — federated sibling atlas.en · Next: B10 — Competency, Assessment & Certification Pathway.es · Domain B: Patient Assessment & Consultation