1. เริ่มจากชนิดของ claim
Frequency/descriptive
Thirty per cent of respondents reported difficulty.
ต้องการ sample และ measurement ที่เป็นตัวแทน
Associative
Sleep duration was associated with recall.
ต้องการการวัดสองตัวแปรและ control ที่เหมาะ แต่ยังไม่รับรอง cause
Causal
Longer sleep improved recall.
ต้องการ comparison ที่ช่วยตัด confounding และ reverse causation
Predictive
The model accurately predicts hospital admission.
ต้องการ out-of-sample performance, calibration และ population ที่เกี่ยวข้อง
Normative
Hospitals should adopt the model.
นอกจาก accuracy ยังต้องดู harms, fairness, cost และ alternatives
Design ที่ดีสำหรับ claim หนึ่งอาจไม่พอสำหรับอีก claim
2. Construct validity: วัดสิ่งที่อ้างจริงหรือไม่
งานกำหนด engagement เป็นจำนวน clicks
ปัญหา:
- click อาจเกิดโดยไม่อ่าน
- ผู้ใช้ที่เข้าใจเร็วอาจ click น้อย
- meaningful participation อาจเกิดนอกระบบ
จำนวน click มี reliability สูงได้แต่ construct validity ต่ำ
อีกตัวอย่าง:
Productivity was measured by hours logged into the system.
hours logged อาจวัด presence มากกว่า output
ก่อนรับ claim ให้ถาม:
- Concept มีหลายมิติหรือไม่
- Measure ครอบคลุมมิติใด
- มี validation ต่อ criterion อื่นหรือไม่
- Measure ทำให้พฤติกรรมเปลี่ยนหรือไม่
3. Sampling และ representativeness
A survey of platform subscribers found that 82% support the new premium feature.
ผลนี้อาจใช้กับ active subscribers แต่ไม่ควรขยายถึง general public หรือ former users
Sampling problems:
- convenience sample
- non-response
- exclusion ของผู้ไม่มี internet
- survivor bias
- sample จาก setting เดียว
Sample ใหญ่ไม่แก้ selection bias หากทุกคนมาจากกลุ่มที่ผิด
A million responses from voluntary website visitors may be less representative than a carefully selected sample of one thousand.
4. Selection bias
App สำรวจเฉพาะผู้ทำ challenge ครบ 100 วัน แล้วรายงาน satisfaction สูง
ผู้ที่ทำครบอาจ:
- พึงพอใจมากกว่า
- มี motivation สูงกว่า
- มีปัญหาน้อยกว่า
- ได้ประโยชน์มากกว่า
ผู้ที่เลิกกลางทางหายจาก sample จึงประเมินผลโดยรวมสูงเกินจริงได้
การถามเฉพาะ completers ไม่ผิดหาก claim คือ “ประสบการณ์ของผู้ทำครบ” แต่ผิดหาก claim คือ “ผู้เริ่มโปรแกรมทุกคนจะพึงพอใจ”
5. Confounding
Ice-cream sales and drownings rise together.
ไม่ควรสรุปว่า ice cream causes drowning เพราะอากาศร้อนอาจเพิ่มทั้งการซื้อไอศกรีมและการว่ายน้ำ
Confounder ต้องสัมพันธ์กับ exposure และ outcome และไม่ใช่เพียงผลตาม causal path
ตัวอย่างการศึกษา tutoring:
Students who chose tutoring scored higher.
Possible confounders:
- prior attainment
- motivation
- parental support
- available study time
Regression adjustment ช่วยเฉพาะ confounders ที่วัดได้และ model เหมาะ ไม่สามารถรับรองว่าไม่มี residual confounding
6. Reverse causation
Employees who felt less stressed exercised more.
Possible directions:
- exercise reduced stress
- lower stress made exercise easier
- third factor เช่น flexible schedules affected both
Cross-sectional data วัดพร้อมกันจึงแยกทิศทางยาก Longitudinal design ช่วย temporal ordering แต่ยังไม่ตัด confounding ทั้งหมด
คำว่า predicts ทางสถิติไม่จำเป็นต้องหมายถึงเกิดก่อนหรือเป็นสาเหตุ หากใช้ใน model contemporaneous
7. Comparison group และ randomisation
Random assignment ช่วยทำให้ known/unknown confounders กระจายใกล้กันโดยเฉลี่ย แต่ต้องตรวจ:
- allocation concealment
- baseline imbalance
- adherence/crossover
- attrition
- blinding ของ outcome assessment
- analysis ตาม assignment หรือ treatment received
Randomised trial ไม่ได้สมบูรณ์โดยอัตโนมัติ และอาจมี external validity จำกัด
Observational study ไม่ได้ไร้ค่า มันอาจเหมาะกับ rare harms, long-term effects หรือ exposure ที่สุ่มไม่ได้ แต่ causal language ต้องระมัดระวัง
8. Missing data และ attrition
Attrition น่ากังวลเมื่อ:
- ต่างกันระหว่างกลุ่ม
- สัมพันธ์กับ outcome
- เหตุผลไม่ถูกบันทึก
- ผู้วิเคราะห์ใช้เฉพาะ completers
ตัวอย่าง:
Satisfaction averaged 9/10 among respondents, but only 35% completed follow-up.
หากผู้ไม่พอใจไม่ตอบ ค่าเฉลี่ยอาจสูงเกินจริง
คำว่า missing at random มีความหมายทางสถิติภายใต้ข้อมูลที่สังเกต ไม่ได้แปลว่าข้อมูลหายแบบสุ่มธรรมดา
9. Precision และ confidence interval
Point estimate:
The treatment reduced symptoms by 8%.
หาก 95% CI คือ −3% ถึง 19% ข้อมูลยังเข้ากันได้กับ harm เล็กน้อย ไม่มีผล หรือ benefit ที่มีขนาดพอสมควร
ข้อความที่เหมาะ:
The estimate favoured treatment, but it was imprecise and did not rule out little or no benefit.
ไม่ควรเขียน:
The treatment had no effect because the result was not significant.
Non-significance ไม่บอก effect เท่ากับศูนย์ ต้องดู interval และ power
10. Statistical significance กับ practical importance
Sample ใหญ่มากอาจทำให้ effect เล็กมาก significant
Average waiting time fell by 12 seconds, p < .001.
ต้องถามว่า 12 วินาทีมีความหมายต่อผู้ป่วยหรือ operation หรือไม่
ตรงกันข้าม effect ที่สำคัญทางปฏิบัติอาจไม่ significant ใน small sample เพราะ imprecision
การประเมินควรดู:
- effect size
- uncertainty
- baseline risk
- costs/harms
- minimal important difference
11. Base rate และ predictive value
สมมติ disease พบ 1 ใน 1,000 และ test มี sensitivity/specificity สูง ผลบวกยังอาจมี false positives จำนวนมากเมื่อเทียบ true positives
คำถามว่า “test accurate 99%” ไม่พอ ต้องรู้:
- prevalence/base rate
- sensitivity
- specificity
- positive predictive value ใน population นี้
Base-rate neglect ทำให้ผู้อ่านตีความผลคัดกรองว่าเป็น diagnosis ที่แน่นอนเกินไป
12. Multiple testing และ exploratory analysis
หากทดสอบ outcomes หรือ subgroups จำนวนมาก บางผลอาจ significant โดย chance
สัญญาณ:
- outcome ไม่ได้ prespecify
- subgroup เล็กหลังเห็นข้อมูล
- รายงานเฉพาะผล positive
- model หลายรูปแต่ไม่กล่าว sensitivity analysis
Exploratory finding มีประโยชน์สร้าง hypothesis แต่ควรถูกเรียกว่า provisional และต้อง confirm ในข้อมูลใหม่
A dramatic result in one small subgroup is not equivalent to a replicated primary outcome.
13. Source credibility กับ evidence relevance
ผู้เชี่ยวชาญที่น่าเชื่อถืออาจให้หลักฐานที่ไม่ตรง claim
A leading surgeon endorses a national education policy.
ชื่อเสียงด้าน surgery ไม่ทำให้เป็นผู้เชี่ยวชาญ education
ในทางกลับกัน source ที่มี conflict of interest ไม่ได้ทำให้ข้อมูลผิดโดยอัตโนมัติ แต่เพิ่มเหตุให้ตรวจ methods, transparency และ independent replication
ควรแยก:
- Expertise
- Direct access to evidence
- Method quality
- Incentives/conflicts
- Corroboration
14. Replication และ robustness
ผลที่แข็งแรงควร:
- ทำซ้ำในข้อมูลใหม่
- คง direction ภายใต้ reasonable analyses
- ไม่ขึ้นกับ outliers เพียงไม่กี่จุด
- มี measurement ที่เทียบได้
- แสดง heterogeneity อย่างโปร่งใส
Failure to replicate ไม่พิสูจน์ทันทีว่าผลแรก fraud อาจต่าง sample, method หรือ power แต่ลด confidence และต้องวิเคราะห์ divergence
15. Worked Example
Claim:
A language-learning app doubles vocabulary growth.
Evidence:
Users who purchased the premium plan learned an average of 40 words, compared with 20 among free users over one month.
Evaluation
- Observational/self-selected groups
- Premium users อาจมี motivation/ability/time ต่าง
- ไม่รู้ measurement reliability
learned อาจเป็น recognition test ไม่ใช่ productive vocabulary- One-month outcome ไม่บอก retention
- Average ratio 40/20 ไม่รับรองว่าทุกคน doubled
Revised claim:
Premium users recorded greater one-month vocabulary gains, although self-selection and uncertainty about long-term retention prevent a causal conclusion.
16. วิธีประเมินหลักฐานทีละขั้น
- ระบุชนิดและขอบเขตของ claim
- ตรวจว่า measure operationalise concept อย่างไร
- ตรวจ sampling และ comparison
- หา selection, confounding และ reverse causation
- ตรวจ missing data และ analysis choices
- ดู effect size กับ uncertainty ไม่ใช่ p-value อย่างเดียว
- พิจารณา alternatives และ replication
- เขียน conclusion ใหม่ให้ตรงกับ evidence strength
17. กับดักที่พบบ่อย
- เชื่อ sample ใหญ่ว่าแก้ bias ทุกชนิด
- สรุป causation จาก self-selection
- ใช้ clicks แทน engagement โดยไม่วิจารณ์ construct
- ตีความ non-significant ว่าไม่มี effect
- มอง significant ว่าสำคัญทางปฏิบัติ
- ลืม base rate
- เชื่อ authority แทน method
- ปฏิเสธงานทั้งหมดเพราะมี limitation หนึ่งจุด
18. Checklist ก่อนตอบ
- Claim ต้องการ inference ชนิดใด
- Measure ตรงกับ construct หรือไม่
- Sample เป็นตัวแทนใคร
- Comparison ลด confounding ได้เพียงใด
- Missing data ทำให้ผลเอียงหรือไม่
- Estimate แม่นและสำคัญทางปฏิบัติหรือไม่
- Alternative explanations ยังเหลืออะไร
- Conclusion ควรลด strength หรือ scope ตรงไหน
Critical evaluation ระดับ C2 ไม่ได้ถามเพียงว่าหลักฐาน “ดีหรือไม่ดี” แต่ระบุอย่างแม่นว่าหลักฐานตอบคำถามใดได้ ตอบได้ด้วยความแน่นอนเท่าใด และคำถามใดยังอยู่นอกขอบเขตของมัน