Hi again,
The first thanks, I’ve been able to keep progressing with the tests, and as of now, I don’t think the problem is in the EC or the VBIOS.
My Bios request topic here -> https://www.badcaps.net/forum/troubl...ur#post3828581
After running MATS with the correct version for AD107, the test completes, detects the GPU correctly, and the memory bus responds, but write errors appear consistently in the same subpartitions and with the same bit patterns. This is typical of a physical VRAM failure, not a thermal flag or firmware issue.
At this point, I’m mostly ruling out:
About MODS
With MODS, the behavior is slightly different and provides useful insight:
FBFALCON_MAILBOX(03) : 0x75880398
...
FALCON_TRACEPC(0) : 0x00003786
FALCON_TRACEPC(1) : 0x0000377f
...
Error Code = 000000000539 (NVRM Generic falcon error)
MODS reports channel/slot 4 fails to initialize entirely
My interpretation so far (without assuming which physical chips are failing yet):
I’m not assuming yet which chips correspond to which errors, I’m just showing the MATS and MODS data. I’m currently verifying in the boardview which slot is actually empty to avoid confusion with a faulty chip.
I’m currently:
If anyone already has a mapping (physical position ↔ memory channel / FBIO) or experience with MODS showing a channel that doesn’t initialize while others show errors, it would be extremely helpful. This could save me from repeatedly removing the heatsink.
Everything points to replacing the failing chips as the next step, but I want to fully understand the MODS behavior first.
Thanks again for the help.
P.D. MATS log:
Found NVIDIA devices :
[ 0] VGA : 0000:01:00.0
RM unavailable - won't test GALIT2 FrameBuffer upper memory
Half-FBPA mask is 0x0, but expected 0x1
mats version 520.249. Testing AD107 with 10 MB of memory starting with 0 MB.
Memory Errors on
Read Error Count: 0
Write Error Count: 272
Unknown Error Count: 0
=== MEMORY ERRORS BY SUBPARTITION ===
SUBPART READ ERRORS WRITE ERRORS UNKNOWN ERRS
------- ----------- ------------ ------------
FBIOA0 0 0 0
FBIOA1 0 104 0
FBIOB0 0 64 0
FBIOB1 0 104 0
Failing Bits:
A032-A063
B000-B063
P.D2:position in boardview
à
The first thanks, I’ve been able to keep progressing with the tests, and as of now, I don’t think the problem is in the EC or the VBIOS.
My Bios request topic here -> https://www.badcaps.net/forum/troubl...ur#post3828581
After running MATS with the correct version for AD107, the test completes, detects the GPU correctly, and the memory bus responds, but write errors appear consistently in the same subpartitions and with the same bit patterns. This is typical of a physical VRAM failure, not a thermal flag or firmware issue.
At this point, I’m mostly ruling out:
- EC flag
- Thermal flag in VBIOS
About MODS
With MODS, the behavior is slightly different and provides useful insight:
- MODS reports that one memory channel/slot fails to initialize entirely, while the others do enter training.
- Additionally, Falcon logs show generic errors and trace entries, indicating that the GPU attempts to initialize the memory channel but fails. For example:
FBFALCON_MAILBOX(03) : 0x75880398
...
FALCON_TRACEPC(0) : 0x00003786
FALCON_TRACEPC(1) : 0x0000377f
...
Error Code = 000000000539 (NVRM Generic falcon error)
MODS reports channel/slot 4 fails to initialize entirely
- This matches what MATS shows:
- MATS runs but reports errors in two specific subpartitions
- MODS shows one channel that doesn’t even train / initialize
My interpretation so far (without assuming which physical chips are failing yet):
- Two VRAM subpartitions are clearly faulty (the ones showing MATS errors)
- One slot is empty (as expected for this SKU), which MODS reports as not initializing
- Falcon logs reinforce that the GPU finds initialization problems on that channel
I’m not assuming yet which chips correspond to which errors, I’m just showing the MATS and MODS data. I’m currently verifying in the boardview which slot is actually empty to avoid confusion with a faulty chip.
I’m currently:
- Cross-checking MATS results with VRAM1–VRAM4 numbering in the boardview
- Physically identifying which chips correspond to the failing subpartitions
- Confirming which slot is truly empty on this FA707NUR SKU
If anyone already has a mapping (physical position ↔ memory channel / FBIO) or experience with MODS showing a channel that doesn’t initialize while others show errors, it would be extremely helpful. This could save me from repeatedly removing the heatsink.
Everything points to replacing the failing chips as the next step, but I want to fully understand the MODS behavior first.
Thanks again for the help.
P.D. MATS log:
Found NVIDIA devices :
[ 0] VGA : 0000:01:00.0
RM unavailable - won't test GALIT2 FrameBuffer upper memory
Half-FBPA mask is 0x0, but expected 0x1
mats version 520.249. Testing AD107 with 10 MB of memory starting with 0 MB.
Memory Errors on
Read Error Count: 0
Write Error Count: 272
Unknown Error Count: 0
=== MEMORY ERRORS BY SUBPARTITION ===
SUBPART READ ERRORS WRITE ERRORS UNKNOWN ERRS
------- ----------- ------------ ------------
FBIOA0 0 0 0
FBIOA1 0 104 0
FBIOB0 0 64 0
FBIOB1 0 104 0
Failing Bits:
A032-A063
B000-B063
P.D2:position in boardview
Comment