Opal problem causing crash but old compiler works [WAS] Re: [Vm-dev] Debugging VM crash, INVALID RECEIVER / a(n) bad class ??
Hi guys, So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly. I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare? I also tried in Pharo 5.0 but same results. Thanks in advance, On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com> :
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine. I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again. The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.* Any clues? createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com
:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via Smalltalk compilerClass: XXX. (I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
Hi Nicolai, Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well. I have updated https://pharo.fogbugz.com/f/cases/13854 Thanks Nicolai!! On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com> :
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com
wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
2015-11-17 14:21 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Hi Nicolai, Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well. I have updated https://pharo.fogbugz.com/f/cases/13854
Thanks Nicolai!!
Hi Mariano, I am glad that it works! I reviewed the fix again (wasn't sure if all changes were needed - but yes they are) and added a slice for issue 13854. If it passes the tests and works in Pharo 5.0, we can create another issue and slice for Pharo 4.0 too. nicolai
On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com
:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess < nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
Thanks Nicolai. Our monkey say it was success :) Do you think Marcus should review it? or you confident enough? Cheers, On Wed, Nov 18, 2015 at 6:43 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 14:21 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Hi Nicolai, Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well. I have updated https://pharo.fogbugz.com/f/cases/13854
Thanks Nicolai!!
Hi Mariano, I am glad that it works! I reviewed the fix again (wasn't sure if all changes were needed - but yes they are) and added a slice for issue 13854. If it passes the tests and works in Pharo 5.0, we can create another issue and slice for Pharo 4.0 too.
nicolai
On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess < nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
2015-11-18 13:53 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Thanks Nicolai. Our monkey say it was success :) Do you think Marcus should review it? or you confident enough?
A review would be good :) How about you, Henrik? This is an interesting bug :-) nicolai
Cheers,
On Wed, Nov 18, 2015 at 6:43 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 14:21 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Hi Nicolai, Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well. I have updated https://pharo.fogbugz.com/f/cases/13854
Thanks Nicolai!!
Hi Mariano, I am glad that it works! I reviewed the fix again (wasn't sure if all changes were needed - but yes they are) and added a slice for issue 13854. If it passes the tests and works in Pharo 5.0, we can create another issue and slice for Pharo 4.0 too.
nicolai
On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com> :
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess < nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
On 18 Nov 2015, at 5:20 , Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-18 13:53 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com <mailto:marianopeck@gmail.com>>: Thanks Nicolai. Our monkey say it was success :) Do you think Marcus should review it? or you confident enough?
A review would be good :)
How about you, Henrik? This is an interesting bug :-)
My initial though was "Why can't we just pop the copied vars from the stack *after* the block creation, instead of before, when they're part of the closures stack depth?", but didn't have the time to actually understand how it works/why that would be wrong... So I'll trust your judgement on this. To bikeshed, I would have preferred stack pop: anAmount to anAmount timesRepeat: [stack pop] though. ;) Cheers, Henry
2015-11-18 13:53 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Thanks Nicolai. Our monkey say it was success :) Do you think Marcus should review it? or you confident enough?
Now there is a bug tracker entry for pharo 4.0 too, and a fix, ready for testing. 17057 <https://pharo.fogbugz.com/f/cases/17057/BackPort-Pharo4-13854-frameSize-calc...> BackPort Pharo4: 13854 frameSize calculated wrongly for #lineSegmentsDo:
Cheers,
On Wed, Nov 18, 2015 at 6:43 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 14:21 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:
Hi Nicolai, Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well. I have updated https://pharo.fogbugz.com/f/cases/13854
Thanks Nicolai!!
Hi Mariano, I am glad that it works! I reviewed the fix again (wasn't sure if all changes were needed - but yes they are) and added a slice for issue 13854. If it passes the tests and works in Pharo 5.0, we can create another issue and slice for Pharo 4.0 too.
nicolai
On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com> :
On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:
2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
Hi guys,
So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.
I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56). Any idea what else can I check/compare?
I also tried in Pharo 5.0 but same results.
Thanks in advance,
OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin... But the attached cs in there does NOT fixes mine.
You need to: switch to old compiler in settings
I was doing that via
Smalltalk compilerClass: XXX.
(I didn't know there was setting).
load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!) switch back to opal recompile all
Yes, I did that.
But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are many changes for the bytecodegenerator.
Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator
Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?
you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all of this code is necessary - and correct) : fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).
I can confirm it's a problem with the number of temp vars defined. As soon as I remove any of the temp vars, it works again.
The method is this (a bit ugly, yes) pasted below. But the key point is: *As soon as I remove ANY tempVar the code starts to work again.*
Any clues?
createZacksAccountingRulesFromTable: t1 | t2 t3 t4 | t2 := FaUserContextInformation current dbAccessor db: 'zacks' getTableAccessorOn: t1 indexOn: nil. t3 := OrderedCollection new. t4 := FaFileTranscript named: 'create-zacks-rules.txt'. FaApplicationDB session inUnitOfWorkDo: [:t5 | | t6 | t6 := FaUserContextInformation current userDisplayName. t2 doWithRowDictionaries: [:t7 | (t7 at: 'label') == FaNullDatum instance ifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}} do: [:t8 | *| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |* t4 crLog: t8 first. t9 := t8 at: 1. t16 := t7 at: 'label'. t18 := t7 at: 'TTM'. t10 := t7 at: 'quuveKeyEquivalent'. t12 := (t7 at: 'zacksKey') asUppercase. t13 := 'ZACKS_' , t12. t14 := 'ZACKS_Q_' , t12. t15 := 'ZACKS_A_' , t12. t11 := (t8 at: 2) , t12. t17 := FaAccountingRule new label: t16; selector: t13 asSymbol; context: 'Zacks' , t9 , 'Override'; argumentTypesSpecification: '{FaProcessorProxy}'; returnTypeSpecification: 'FaDatedFnSeries'; isReturnValueConstant: true; action: (self getZacksDataAtScriptForCache: t11 forZacksKey: t11 fromSet: t9); comment: (nil ifNil: [t16]); isSensitive: true; definedBy: t6; ownedBy: t6; lastEditBy: t6; yourself. t5 register: t17. t3 add: 'Successfully added ' , t17 selector , '/' , t17 context. t4 crLog: t3 last]]]]. ^ t3
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck < marianopeck@gmail.com> wrote:
On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess < nicolaihess@gmail.com> wrote:
Hi Mariano, if you know the method that may cause this crash can you check if it works if you recompile this method with the old compiler, maybe this issue is responsible - wrong stack frame size : 13854 <https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lin...> frameSize calculated wrongly for #lineSegmentsDo:
Hi Nicolai,
Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling). However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash. So it's definitively something related to Opal compilation, and probably, related to closures compilation. I will see if I find other opal issues opened in 4.0.
Thanks!
2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck < marianopeck@gmail.com>:
Hi guys,
I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.
I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure* *0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3721e1f0 is not a context* *0x371c34b4 is not a context*
And another crash:
*Program received signal SIGSEGV, Segmentation fault.* *0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:40444* *40444 && (((longAt(fieldOop)) & MarkBit) != 0)) {* *(gdb) call printAllStacks()* *Process 0x30e228c4 priority 40* *0xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class* *0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class* *0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class* *0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class* *0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space*
*(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 48626* *0x3722b750 is not a context* *[New Thread 0x1b43 of process 5886]* *[New Thread 0x1d2f of process 5886]* *[New Thread 0x1f07 of process 5886]* *[New Thread 0x1e07 of process 5886]*
From what I can see in *#printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessage* It looks like if the memory address is not an OOP, nor forwarding pointer nor...
It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)
As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.
BTW.... in the output of the gdb I see lots of printings like:
*(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421* *(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416*
Any pointer is appreciated.
Thanks!
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
-- Mariano http://marianopeck.wordpress.com
On Sat, Nov 21, 2015 at 4:31 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Now there is a bug tracker entry for pharo 4.0 too, and a fix, ready for testing. 17057 BackPort Pharo4: 13854 frameSize calculated wrongly for #lineSegmentsDo:
Minor thought: The earlier comment on the original issue considered this a significant change that carried risk to implement prior to Pharo 4 release, and I think the same should be considered for backports. Perhaps see how the dust settles for a bit on Pharo 5, before integrating into Pharo 4. @Mariano, can you hang on and for a while load the slice manually or through a startup loader? cheers -ben
On Sat, Nov 21, 2015 at 10:36 AM, Ben Coman <btc@openinworld.com> wrote:
On Sat, Nov 21, 2015 at 4:31 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:
Now there is a bug tracker entry for pharo 4.0 too, and a fix, ready for testing. 17057 BackPort Pharo4: 13854 frameSize calculated wrongly for #lineSegmentsDo:
Minor thought: The earlier comment on the original issue considered this a significant change that carried risk to implement prior to Pharo 4 release, and I think the same should be considered for backports. Perhaps see how the dust settles for a bit on Pharo 5, before integrating into Pharo 4. @Mariano, can you hang on and for a while load the slice manually or through a startup loader?
Sure, in fact it's already part of my image setup building :) So.... no need from me to be integrated since I can get the slice. And agree..maybe some testing in 5.0 may tell us how safe is to backport to 4.0. The good news is that it seems that also Marcus took a look too. Best, -- Mariano http://marianopeck.wordpress.com
participants (4)
-
Ben Coman -
Henrik Johansen -
Mariano Martinez Peck -
Nicolai Hess