Restore macOS Intel - #565
Conversation
|
@mscheltienne a different run is failing here, it's flaky. I'll keep looking to see if I can make it more reliable somehow (maybe get Claude to see if it has ideas...) |
|
@mscheltienne I threw Claude Opus 4.8 at the timeout problem with the suggestion of using |
I had this suggestion several times, it's a workaround instead of digging through the root cause which is why I haven't implemented it yet (it potentially hides failures that we don't have figure out yet). In this case, I tend to push Claude to dig for the root cause by making hypothesis then creating debug scripts in e,g, A. Over-count (4 epochs, expected 3) —
|
|
But you are also using |
|
So yes.. Anyway, I think the fix above and one more improvement baking now should actually yield stable CIs, if not I'll re-introduce the |
|
I'll trigger the CIs 5/6 times to check for remaining flakiness. For now: 2 / 2. |
|
Ahh yes I didn't think of segfault. But maybe pytest-xdist plus rerun would also work for this case (can't remember but it might) to stay more targeted Hopefully it won't be needed at all! 🤞 |
|
I definitely remember looking into that.. but can't remember what was the issue with |
|
Coming back green again, time to remove the |
|
The retry-step would be necessary at least until sccn/liblsl#289 is reviewed and merged - Claude flags that this error is currently caught by the retry-step and thus "safe". |
|
But at least, it looks like flakiness in our python side is down! |
|
Yeah once it looks good to go, just ask Fable about the possibility of replacing the action-retry with pytest-xdist (even with n=1, which seems necessary for all this network business) plus pytest-rerunfailures... I thought I saw that working in some repo recently but maybe not! But if it does work, would be nice to slot it in here as well. (Or I can do it if you need to move on to other stuff.) |
|
Go for it, thanks 🙏 |
|
Looks like we're doing well, Fable had more ideas, but I'm going to have it pursue (1) and (2) because they seem low-risk: IdeasYes — several, and the vendored liblsl source lets me ground them in the actual crash path. Recall the chain: close_stream() breaks the connection → receiver thread calls try_recover_from_error() → resolver launches resolve waves ("re-connecting…" logs) → lsl_destroy_inlet races with an in-flight resolve attempt → shared_from_this() throws bad_weak_ptr → unwind skips the thread join → std::terminate. Ordered by leverage:
My recommendation: 1 and 2 are complementary and both Python-side-available today — 1 removes the provocation (recovery engaging at teardown), 2 removes the machinery (resolve attempts existing at all). I'd try 1 first since it's least invasive and arguably a correctness cleanup on its own, keep the xdist rerun as the safety net, and treat 3 as the durable fix. So I'll push these, and then I think we should merge assuming it comes back green again! |
|
For the proposal (3), as soon as 1.18 stable is released (I think we are now at the 4th RC release), I will update the |
|
Can you also delete the now unused retry action? |
|
I think it's still used for doc building |
|
Looks like doc build has been stable, I'll remove the retry for it and the action |
|
Okay good to merge I think @mscheltienne ! |
|
Thanks a lot! |
|
Hopefully things stay stable now! |
Hopefully this is all that's needed! 🤞