On 19 June 2018 at 22:55, Cyril Ferlicot D. <cyril.ferlicot@gmail.com> wrote:
Hi,
Since months now there are a lot of random failure on the CI making it hard to work.
There is different kind of failures: - Network problems - Failing tests - Incomprehensible problems
Now I don't see much failure due to Network. I suppose the Inria infrastructure improved.
Failing tests were corrected those past months and we see less and less of them.
Now the big problem are the incomprehensible crashes such as "The workspace was not found" or "FileDoesNotExistException" or "pharo-vm/ is already present".
We just found the problem :)
During the validation of the Bootstrap multiple tests are launched on OSX/Windows/linux in parallel. Each task is on a different slave of the Jenkins. But, apparently we discovered that two slaves could have the same disk. Usually it does not cause any trouble since a job is only run by one slave. But in this particular case, two slaves can be used by the same job and mess with the resources of each other.
That sort of outside-the-box confounding factor is difficult and frustrating to track down. Great work guys. cheers -ben