Hi Nicolai,Thanks for sharing an updated cs. I can confirm that this version DOES work and fixes my VM crash as well.��I have updated��https://pharo.fogbugz.com/f/cases/13854Thanks Nicolai!!��
--On Tue, Nov 17, 2015 at 6:36 AM, Nicolai Hess <nicolaihess@gmail.com> wrote:2015-11-17 1:35 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:On Mon, Nov 16, 2015 at 8:41 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:2015-11-16 22:16 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:On Mon, Nov 16, 2015 at 5:37 PM, Mariano Martinez Peck <marianopeck@gmail.com> wrote:Hi guys,So I found out the exact method that is causing the VM crash. If I compile the method with Opal, it crashes the VM when I execute the code that use that method. If I compile it with old compiler, the code does work correctly.��I tried comparing both compiled methods compiled from both Compilers and I cannot see real differences. They both seem to have similar (same?) bytecodes, literals, decompiled string, etc. The only difference I see is in #frameSize (old compiler one is 16 while Opal one is 56).��Any idea what else can I check/compare?��I also tried in Pharo 5.0 but same results.��Thanks in advance,��OK, it seems my issue may be related to: https://pharo.fogbugz.com/f/cases/13854/frameSize-calculated-wrongly-for-lineSegmentsDoBut the attached cs in there does NOT fixes mine.You need to:switch to old compiler in settingsI was doing that via��Smalltalk compilerClass: XXX.(I didn't know there was setting).��load change set fix_closure_stack_frame_size_computation.1.cs (not fix_closure_stack_frame_size_computation.2.cs!)switch back to opalrecompile allYes, I did that.����But the change set is old, and does not work for recent images (not even for pharo 4.0 release), because there are manychanges for the bytecodegenerator.Uhhhh. Yes, I saw some. For example, i changed from IRAbstractBytecodeGenerator to IRBytecodeGenerator��Do you think we can re-create your fix with latest 4.0 or 5.0 so that I can see if the issue was the same?you can try this one, (I just changed all the code that makes this run for pharo 4, but I need more time to check if all ofthis code is necessary -�� and correct) :fix_closure_stack_size_4.0.cs (tested with Pharo 4.0 40624).����
��I can confirm it's a problem with the number of temp vars defined.�� As soon as I remove any of the temp vars, it works again.The method is this (a bit ugly, yes) pasted below.But the key point is: As soon as I remove ANY tempVar����the code starts to work again.Any clues?createZacksAccountingRulesFromTable: t1��| t2 t3 t4 |t2 := FaUserContextInformation current dbAccessordb: 'zacks'getTableAccessorOn: t1indexOn: nil.t3 := OrderedCollection new.t4 := FaFileTranscript named: 'create-zacks-rules.txt'.FaApplicationDB sessioninUnitOfWorkDo: [:t5 |��| t6 |t6 := FaUserContextInformation current userDisplayName.t2doWithRowDictionaries: [:t7 | (t7 at: 'label')== FaNullDatum instanceifFalse: [{{'Annual'. 'ZACKS_A_'}. {'Quarterly'. 'ZACKS_Q_'}}do: [:t8 |��| t9 t10 t11 t12 t13 t14 t15 t16 t17 t18 |t4 crLog: t8 first.t9 := t8 at: 1.t16 := t7 at: 'label'.t18 := t7 at: 'TTM'.t10 := t7 at: 'quuveKeyEquivalent'.t12 := (t7 at: 'zacksKey') asUppercase.t13 := 'ZACKS_' , t12.t14 := 'ZACKS_Q_' , t12.t15 := 'ZACKS_A_' , t12.t11 := (t8 at: 2), t12.t17 := FaAccountingRule new label: t16;selector: t13 asSymbol;context: 'Zacks' , t9 , 'Override';argumentTypesSpecification: '{FaProcessorProxy}';returnTypeSpecification: 'FaDatedFnSeries';isReturnValueConstant: true;action: (selfgetZacksDataAtScriptForCache: t11forZacksKey: t11fromSet: t9);comment: (nilifNil: [t16]);isSensitive: true;definedBy: t6;ownedBy: t6;lastEditBy: t6;yourself.t5 register: t17.t3 add: 'Successfully added ' , t17 selector , '/' , t17 context.t4 crLog: t3 last]]]].^ t3��On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck <marianopeck@gmail.com> wrote:--On Mon, Nov 16, 2015 at 2:48 PM, Mariano Martinez Peck <marianopeck@gmail.com> wrote:On Mon, Nov 16, 2015 at 1:07 PM, Nicolai Hess <nicolaihess@gmail.com> wrote:��maybe this issue is responsible - wrong stack frame size :Hi Mariano, if you know the method that may cause this crash can you checkif it works if you recompile this method with the old compiler,
13854 frameSize calculated wrongly for #lineSegmentsDo:Hi Nicolai,Thanks for the pointer. I tried with your fix and I still get the same results (even after recompiling).However...if I swap back to old Compiler rather than OpalCompiler and I recompile everything, then I do not have anymore the crash.��So it's definitively something related to Opal compilation, and probably, related to closures compilation.��I will see if I find other opal issues opened in 4.0.Thanks!��2015-11-16 17:00 GMT+01:00 Mariano Martinez Peck <marianopeck@gmail.com>:��Hi guys,I am debugging a Pharo VM crash I am having and I cannot figure out what it is exactly. I suspect it might be related to block closures compilation but I am not sure. I have a reproducible crash test under Pharo 4.0 and OSX.��I built a VM in debug mode and I run it via gdb. This is the kind of info I am able to see:Program received signal SIGSEGV, Segmentation fault.0x000ab66b in updatePointersInRangeFromto (memStart=924623948, memEnd=928201420) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:4044440444 && (((longAt(fieldOop)) & MarkBit) != 0)) {(gdb) call printAllStacks()Process 0x30e228c4 priority 400xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class0xbffb2e30 M BlockClosure(FaMemoryStoreSession)>register: 0x371d0968: a(n) BlockClosure0xbffb2e4c M INVALID RECEIVER>register: 0x371cfd54: a(n) bad class(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 486260x3721e1f0 is not a context0x371c34b4 is not a contextAnd another crash:Program received signal SIGSEGV, Segmentation fault.0x000ab66b in updatePointersInRangeFromto (memStart=924677864, memEnd=928262300) at /Users/mariano/Pharo/git/pharo-vm/src/vm/gcc3x-cointerp.c:4044440444 && (((longAt(fieldOop)) & MarkBit) != 0)) {(gdb) call printAllStacks()Process 0x30e228c4 priority 400xbffb2de0 M INVALID RECEIVER>initialize 0x3722b9a0: a(n) bad class0xbffb2df8 M FaAction class(Behavior)>new 0x20ebb104: a(n) FaAction class0xbffb2e10 M FaAction class>block: 0x20ebb104: a(n) FaAction class0xbffb2e30 M INVALID RECEIVER>register: 0x371ddee0: a(n) bad class0xbffb2e4c M INVALID RECEIVER>register: 0x371dd328 is in old space(callerContextOrNil == (nilObject())) || (isContext(callerContextOrNil)) 486260x3722b750 is not a context[New Thread 0x1b43 of process 5886][New Thread 0x1d2f of process 5886][New Thread 0x1f07 of process 5886][New Thread 0x1e07 of process 5886]From what I can see in #printActivationNameFor: aMethod receiver: anObject isBlock: isBlock firstTemporary: maybeMessageIt looks like if the memory address is not an OOP, nor forwarding pointer nor...It also seems like the above stack I can get is not the real cause but a side effect that makes GC to crash (#updatePointersInRangeFromto)As said, I have a way to reproduce the crash, and I already have the VM compiled in debug and gdb running. I can also attach the full output of the gdb stack.��BTW.... in the output of the gdb I see lots of printings like:(numStack + ReceiverIndex) < (lengthOf(theContext)) 45421(ReceiverIndex + contextSize) < (lengthOfbaseHeaderformat(oop, header2, fmt)) 40416Any pointer is appreciated.Thanks!--
--
--
--
--