I had an error in a Raid 1 volume with two disks split between two enclosures that was showing an i/o error on the primary disk in Enclosure 1-disk in slot A. I moved the problem disk into the same enclosure as the healthy, still error-ed. I followed the guide on confirming the errors and remediating by using new cables, changing ports, labeling and switching slots. After doing this softraid began showing i/o errors in both slots. Somewhere along the way it also decided to split the volume into two identical volumes, each linked to an individual disk. Ive since purchased an additional Thunderbay 4, so I moved both disks into that and, voila, no i/o errors but the volume/s are still showing as degraded and split.
I'm very new to Softraid, after having almost two decades of issue/error free storage protocols using hardware raid and apple raid, and am finding it wildly frustrating. I've sent emails to support that take multiple days to receive an answer, spent hours delving into faq's, forum posts, etc trying to figure out what the correct protocol is to rebuild the raid with no clear answers (or that I'm able to understand is the answer to my issue-PEBKAC).
There's answers saying Unmount/power down/reconnect- Nope
There's answers saying Rebuild- Not selectable for me.
Answers saying create a new volume- Not possible in the window.
Answers saying Delete/Erase and add back- I cant find a clear distinction between the two. SR app about both- "this deletes data" Does one prevent a rewrite of 16tb? Do you have a vocabulary section anywhere?
Answers saying remove secondary disk and add it back- Followed this one because it appears the least destructive. Nope. Theres not enough space in the available volumes.
Is there a path that rebuilds the volume without having to rewrite 16tb of data? Should I plan on becoming a power-user just to create/maintain a stable raid-1 volume?
Admittedly, the time I've spent trying to diagnose problems I could have wiped the drives, switched to apple raid, and been close to finished with this 30+tb migration.
You were trying to get support via OWC? I am sorry to hear the slow responses, I will investigate.
When the primary of a mirror is unavailable, the mirror "fails over", meaning the secondary is promoted to primary. If the ex primary is connected in the same boot cycle, SoftRAID can put humpty dupty back together again. If there was a restart in between, or another disk removal, it cannot, as there is no way to know for certain what has changed between the two.
Here is what you need to do:
you need to physically look at the two volumes, if you have written any data to either, to determine which is your actual primary.
In Finder, rename that volume (so there is no confusion).
In SoftRAID disable safeguard (select the volume and select disable safeguard from volumes menu)
Delete that volume.
To the remaining volume:
Set Optimization to workstation (else rebuilding will be slow)
Normally you need to "remove missing secondary disk", but not here, so lets skip that step.
Now, "Add disk" and select the removed disk. The volume will be usable immediately and rebuild in the background.
Because of all the events, yes you need to do it this way. But the rebuild is painless and just goes on its way while you work.
Separate enclosures always risks this kind of behavior, as one can be powered up and the other has 15 seconds to show up (There is a time out Mirror preferences in settings to increase this up to two minutes, but it means your volume won't mount for that two minutes when a disk is missing.)
Any other advice you get is incorrect. these are the correct steps.
the IO error could have been hardware, or coudl have been a directory issue, SoftRAID reports ALL disk io reports, regardless of cause. The driver only knows the disk could not read or write. A read error can occur if a corrupted directory tries to read from a non existent location, for example.
Hope this helps!
Appreciate the run through! Also, apologies if this was the wrong section to post this in. I only scrolled fuuurther down and saw the rest of the categories till much later.
It appears I ran through gamut of things that will not get humpty dumpty back together again.
In your experience, is splitting a raid volume between two enclosures a worse practice? Even if you don't mind +1 minute mounts? Or is the likelihood of encountering this issue again that much higher that it negates any mishaps with a single enclosure taking out an array? Asking because I'm looking at switching from my raid-1 setups to raid-10 and was hoping to sneak a bit more throughput. Hesitating now though due to the historical stability I've grown accustomed to.
@bph3647
Thunderbolt can be reliable enough, absolutely. It all depends on what you have around. Is your computer moving? cats? etc.
Get a voltage regulator, however, for better stability if you do not have one It also lets your electronics live longer.

