a11y.skipToMainContent
1

AI รุ่น Benchmark Smission

ลอง เปรียบ เทียบ ม้านั่ง ที่ มี คะแนน ข้าม ผ่าน โมเดล เอ ไอ — เอ็ม แอล ยู, ฮิว แมน เยล, จี เอส เอ็ม 8 เค, และ อีก หลาย อย่าง — เพื่อ จะ หา แบบ จําลอง ที่ ดี ที่ สุด สําหรับ การ ใช้ ของ คุณ.

Benchmark scores from published leaderboards (2025). Higher is better for all metrics.
ModelMMLUHumanEvalGSM8KHellaSwagCost/1M
Click column headers to sort. Scores are approximate and may vary by evaluation method.
Was this tool helpful?
Send output to:
Advertisement

How to use AI Model Benchmark Comparison

  1. เลือกรุ่นที่จะเปรียบเทียบจากรายการ.
  2. ดูคะแนนจากหลายประเภท.
  3. เรียงลําดับตามม้านั่งใด ๆ เพื่อหาต้นแบบที่ดีที่สุดสําหรับงานของคุณ.

อะไรคือ?

โมเดล AI ที่โดดเด่นกว่างานอื่น เครื่องมือนี้ช่วยให้คุณเปรียบเทียบคะแนนมาตรฐานผ่าน GPT-4o, Claude Sonnet, GEmini Pro, Lamia 3, Mistrl, และแบบจําลองอื่น ๆ บนการทดสอบมาตรฐานเช่น MMLU (ความรู้เชิงพาณิชย์), อภิสิทธิ์ (ค.ศ. ใช้การเปรียบเทียบเหล่านี้เพื่อเลือกรุ่นที่เหมาะสม สําหรับภาระงานเฉพาะของคุณ ไม่ว่าจะเป็นรหัส, คณิตศาสตร์, การเขียนสร้างสรรค์ หรือความรู้ทั่วไป

Advertisement

FAQ

ม้านั่งพวกนี้วัดอะไร?
MMLU ทดสอบความรู้ทั่วไปใน 57 เรื่อง มนุษยธรรมวัดรหัสรุ่น GSM8K สอบเหตุผลคณิตศาสตร์ เฮลาสวาเวนประเมินสามัญสํานึก.
คะแนนสูงสุดดีกว่าเสมอ?
สําหรับม้านั่งส่วนใหญ่ ใช่ แต่การแสดงในโลกแห่งความจริง ขึ้นอยู่กับกรณีเฉพาะของคุณ รูปแบบการกระตุ้น และความต้องการความล่าช้า.

Related tools

Author

OH
Omar Hassan"The Number Cruncher"

Engineer & Unit Conversion Specialist

Omar is a mechanical engineer by training and a unit-conversion enthusiast by passion. He has built calibration systems for aerospace and automotive manufacturers and knows firsthand how a single decimal error can cost millions in rework. His mission is to make every conversion instant, accurate, and accessible to everyone, whether they are a student, tradesperson, or practicing engineer, with no advanced degree required.

Advertisement