microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Head-to-head
Compare recorded specifications and inspect source links for both devices. Speed comparisons appear only when workload, runtime version, quantization, context and batch match.
Device profile
Compared here against NVIDIA GeForce RTX 4090 24GB.
Matched rows below
128 GB
60 W
Not scored
Device profile
32B-class Q4 full-GPU at 70+ tok/s — Qwen 3 32B, Qwen 2.5 Coder 32B, QwQ 32B all fit comfortably
Matched rows below
24 GB
450 W
Not scored
Speed comparisons require matching runtime settings. A missing score means insufficient comparative evidence.
Apple M3 Max (40c GPU, 128GB) wins VRAM capacity.
Apple M3 Max (40c GPU, 128GB) has the lower published power rating.
| Metric | Apple M3 Max (40c GPU, 128GB) | NVIDIA GeForce RTX 4090 24GB |
|---|---|---|
| MyAI rating | Not scored | Not scored |
| Speed comparison | See matched rows below | See matched rows below |
| VRAM | 128 GB | 24 GB |
| TDP | 60 W | 450 W |
| MSRP | $4.0k | $1.6k |
| Workloads | 2 | 12 |
This shortlist cross-references the external OpenEvals snapshot with an approximate Q4-class VRAM estimate. Use it to sanity-check whether the hardware you are comparing can host strong open models, not just benchmark toy workloads.
microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Qwen/Qwen2-72B
73B params · est. 43.6 GB Q4
microsoft/Phi-3.5-mini-instruct
3.8B params · est. 2.3 GB Q4
internlm/internlm2_5-7b-chat
7.7B params · est. 4.6 GB Q4