Pharo-dev
By thread
pharo-dev@lists.pharo.org
By month
Messages by month
- ----- 2026 -----
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2008 -----
- December
- November
- October
- September
- August
- July
- June
- May
- 3 participants
- 144622 messages
Re: [Pharo-project] OpenGL in Pharo?
by Serge Stinckwich
On Fri, Nov 18, 2011 at 9:54 AM, Igor Stasenko <siguctua(a)gmail.com> wrote:
> 2011/11/18 Levente Uzonyi <leves(a)elte.hu>:
>> On Thu, 17 Nov 2011, Javier Pimás wrote:
>>
>>> We are advancing on making OpenGL work with NativeBoost right now. If you
>>> have to write an app
>>> that uses OpenGL now, I would strongly recommend you to use NBOpenGL. It's
>>
>> I see the hype about NBOpenGL, but I don't see how is it better than the FFI
>> based OpenGL implementation. Can someone shed some light on it?
>>
>
> The marshalling code is faster. It also deals nice with per-context
> feature availability (depending on context you created, some functions
> may be availabale, some not).
> And it contains a lot of extensions (over 3000 OpenGL functions), up
> to version 3.1 , because all functions are automatically generated
> from specs taken from
> official source.
> A future versions of OpenGL is easy to use: you just regenerate
> sources from specs.
Looks great Igor !
I'm a bit curious, how this generation is automatically done ?
Could you use the same technique for others C librairies ?
Regards,
--
Serge Stinckwich
UMI UMMISCO 209 (IRD/UPMC), Hanoi, Vietnam
Matsuno Laboratory, Kyoto University, Japan (until 12/2011)
http://www.mechatronics.me.kyoto-u.ac.jp/
Every DSL ends up being Smalltalk
http://doesnotunderstand.org/
Nov. 18, 2011
Re: [Pharo-project] WideString performance
by Dennis Schetinin
Thank you very much, I'll play with and surely benefit from your
suggestions.
2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
> 2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> > Also note that the more mainXtream Xtreams at
> > http://www.squeaksource.com/Xtreams/ does not perform that bad...
> >
> > This ones gets a score of 3.0 seconds with a regular MultiByteFileStream
> >
> > nextRow
> > | row char line nextValue noValue |
> > line := stream nextLine reading.
> > row := OrderedCollection new.
> > nextValue := (WideString new: 32) writing.
> > noValue := true.
> >
> > [[[(char := line get) isSeparator] whileTrue.
> > noValue := false.
> > char = $" ifTrue:
> > [[nextValue write: (line ending: $") rest.
> > (char := line get) = $"] whileTrue:
> > [nextValue put: $"].
> > [char = separator] whileFalse: [char := line get]].
> > [char = separator]
> > whileFalse:
> > [nextValue put: char.
> > char := line get].
> > row add: nextValue conclusion.
> > nextValue := (WideString new: 32) writing] repeat]
> > on: Incomplete do: [:exc | ].
> >
> > noValue ifFalse: [row add: nextValue conclusion].
> > ^row
> >
> >
> > with a XTFileReadStream, just rewrite this:
> >
> > MessageTally spyOn: [| stream separator |
> > separator := $;.
> > stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
> > 'Data.csv') reading encoding: #utf8) "rest reading".
> > (stream ending: Character cr) slicing collect: [:line |
> > | row char nextValue noValue |
> > row := OrderedCollection new.
> > nextValue := (WideString new: 32) writing.
> > noValue := true.
> >
> > [[[(char := line get) isSeparator] whileTrue.
> > noValue := false.
> > char = $" ifTrue:
> > [[nextValue write: (line ending: $") rest.
> > (char := line get) = $"] whileTrue:
> > [nextValue put: $"].
> > [char = separator] whileFalse: [char := line
> get]].
> > [char = separator]
> > whileFalse:
> > [nextValue put: char.
> > char := line get].
> > row add: nextValue conclusion.
> > nextValue := (WideString new: 32) writing] repeat]
> > on: Incomplete do: [:exc | ].
> >
> > noValue ifFalse: [row add: nextValue conclusion].
> > row]]
> >
> > For some reason, this was super slow (>20s), though my previous
> > attempt without slicing was around 3.0s...
>
> Ah Ah, replacing contentsSpecies with ^WideString in
> XTEncodeRead/WriteStream took it back to 3.2s.
>
> Nicolas
>
> > (so you noticed the commented "rest reading" whose purpose is to load
> > the whole file in memory)
> >
> > The major advantage is that you can just replace #collect: with
> > #collecting: and get a stream of rows. Lazy is cool.
> >
> > Nicolas
> >
> > 2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> >> 2011/11/17 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> >>> 2011/11/17 Dennis Schetinin <chaetal(a)gmail.com>:
> >>>> So, the answer for the question "Is there a simple way to improve the
> >>>> performance?" would be "No", right?
> >>>>
> >>>
> >>> No, the bottle neck is CSVParser as Henrik and Levente already noticed.
> >>>
> >>> My machine executes your code in 21.7s
> >>>
> >>> So I decided to play with SqueaXTream
> http://www.squeaksource.com/XTream/
> >>> and inlined row decoding in a single method:
> >>>
> >>> CSVParser>>nextRow
> >>> | row char line nextValue |
> >>> line := stream nextLine readXtream.
> >>> row := OrderedCollection new.
> >>> nextValue := (WideString new: 32) writeXtream.
> >>> line endOfStreamAction: [^row add: (nextValue contents);
> yourself].
> >>> [[(char := line next) isSeparator] whileTrue.
> >>> char = $" ifTrue:
> >>> [[nextValue nextPutAll: (stream upTo: $").
> >>
> >> Err, it's line upTo: $"
> >>
> >>> line peekFor: $"] whileTrue:
> >>> [nextValue nextPut: $"].
> >>> [(char := line next) = separator] whileFalse].
> >>> [char = separator]
> >>> whileFalse:
> >>> [nextValue nextPut: char.
> >>> char := line next].
> >>> row add: nextValue contents.
> >>> nextValue reset] repeat
> >>>
> >>> This reduces execution time to 2.4 seconds
> >>>
> >>> I also tried to use SqueaXTream for the file itself (if you don't mind
> >>> the ugly invocation);
> >>> But it does not gain much because you file is essentially made of Wide
> chars...
> >>> It would help only if you'd have more ASCII fields...
> >>>
> >>> MessageTally spyOn: [| stream |
> >>> stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
> >>> 'Data.csv') readXtream binary buffered ~> UTF8Decoder) buffered.
> >>> [(CSVParser on: stream)
> >>> useDelimiter: $;;
> >>> rows ] ensure: [stream close]].
> >>>
> >>> This reduces execution time to 1.9 seconds though
> >>>
> >>> So I wouldn't say there's not plenty of room for performance ;)
> >>>
> >>> Nicolas
> >>>
> >>>> 2011/11/16 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
> >>>>>
> >>>>> I think that there are several issues:
> >>>>>
> >>>>> 1) Character with codePoint < 256 are unique, but Wider are not.
> >>>>> The characters are not stored in the WideString (which instead stores
> >>>>> the code points), but any enumeration using at: rather than wordAt:
> >>>>> will create an overhead. I presume that it would be better to have a
> >>>>> #codePointAt: defined as #byteAt: in ByteString and #wordAt: in
> >>>>> WideString, so we could still define generic methods in String.
> >>>>> 2) A lot of ByteString methods are optimized with ad hoc optional
> >>>>> primitives, but very few WideString methods are.
> >>>>> 3) Some of these optimized operations work with 256 element masks and
> >>>>> cannot be generalized to WideString that easily
> >>>>> 4) Some Stream operations that copy one Collection to the stream
> >>>>> buffer or vice versa are also not optimized for mixed byte/word
> >>>>> collections
> >>>>> 5) UTF8 is optimized for ASCII anyway... And Squeak/Pharo converters
> >>>>> also are (because they can use more primitives in this case)
> >>>>> 6) Pharo ripped #squeakToUtf8 and #utf8ToSqueak which were maybe
> >>>>> questionable from quality POV, but were not replaced from efficiency
> >>>>> POV.
> >>>>>
> >>>>> I would say that a new VM with immediate Character values might help,
> >>>>> but that won't be enough.
> >>>>>
> >>>>> Nicolas
> >>>>>
> >>>>> 2011/11/16 Sven Van Caekenberghe <sven(a)beta9.be>:
> >>>>> > Dennis,
> >>>>> >
> >>>>> > On 16 Nov 2011, at 19:30, Dennis Schetinin wrote:
> >>>>> >
> >>>>> >> Hi!
> >>>>> >>
> >>>>> >> Parsing a UTF-8 CSV file with Russian symbols (using CSVParser), I
> >>>>> >> found it really slow. It takes about 5 seconds to read a 53Kb
> file (136
> >>>>> >> lines) and several minutes to parse a 3.4Mb file (about 8K lines).
> >>>>> >>
> >>>>> >> Profiler shows that nearly all the time is eaten by WideString >>
> >>>>> >> at:put: which is invoked from WriteStream>>nextPut:. If I'm not
> mistaken,
> >>>>> >> this means that <primitive: 66> there does not work in this caseâ¦
> So, to be
> >>>>> >> short: is there a simple way to improve the performance? TIA
> >>>>> >>
> >>>>> >> --
> >>>>> >> Dennis Schetinin
> >>>>> >
> >>>>> > I am interested in this problem.
> >>>>> >
> >>>>> > Would it be possible to provide a test file (not that I can read
> >>>>> > Russian) and your test code ?
> >>>>> >
> >>>>> > I would like to try this myself.
> >>>>> >
> >>>>> > Sven
> >>>>> >
> >>>>>
> >>>>
> >>>>
> >>>>
> >>>> --
> >>>> Dennis Schetinin
> >>>>
> >>>
> >>
> >
>
>
--
Dennis Schetinin
Nov. 18, 2011
Re: [Pharo-project] OpenGL in Pharo?
by Igor Stasenko
2011/11/18 Levente Uzonyi <leves(a)elte.hu>:
> On Thu, 17 Nov 2011, Javier Pimás wrote:
>
>> We are advancing on making OpenGL work with NativeBoost right now. If you
>> have to write an app
>> that uses OpenGL now, I would strongly recommend you to use NBOpenGL. It's
>
> I see the hype about NBOpenGL, but I don't see how is it better than the FFI
> based OpenGL implementation. Can someone shed some light on it?
>
The marshalling code is faster. It also deals nice with per-context
feature availability (depending on context you created, some functions
may be availabale, some not).
And it contains a lot of extensions (over 3000 OpenGL functions), up
to version 3.1 , because all functions are automatically generated
from specs taken from
official source.
A future versions of OpenGL is easy to use: you just regenerate
sources from specs.
I also plan to integrate it with JIT, so when you invoking a method
with native code, it will just run this code without extra checks and
popping out of 'jit' mode.
>
> Levente
--
Best regards,
Igor Stasenko.
Nov. 18, 2011
Re: [Pharo-project] WideString performance
by Nicolas Cellier
2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> Also note that the more mainXtream Xtreams at
> http://www.squeaksource.com/Xtreams/ does not perform that bad...
>
> This ones gets a score of 3.0 seconds with a regular MultiByteFileStream
>
> nextRow
> Â Â Â Â | row char line nextValue noValue |
> Â Â Â Â line := stream nextLine reading.
> Â Â Â Â row := OrderedCollection new.
> Â Â Â Â nextValue := (WideString new: 32) writing.
> Â Â Â Â noValue := true.
>
> Â Â Â Â [[[(char := line get) isSeparator] whileTrue.
> Â Â Â Â noValue := false.
> Â Â Â Â char = $" ifTrue:
> Â Â Â Â Â Â Â Â [[nextValue write: (line ending: $") rest.
> Â Â Â Â Â Â Â Â (char := line get) = $"] whileTrue:
> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue put: $"].
> Â Â Â Â Â Â Â Â [char = separator] whileFalse: [char := line get]].
> Â Â Â Â [char = separator]
> Â Â Â Â Â Â Â Â whileFalse:
> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue put: char.
> Â Â Â Â Â Â Â Â Â Â Â Â char := line get].
> Â Â Â Â row add: nextValue conclusion.
> Â Â Â Â nextValue := (WideString new: 32) writing] repeat]
> Â Â Â Â Â Â Â Â on: Incomplete do: [:exc | ].
>
> Â Â Â Â noValue ifFalse: [row add: nextValue conclusion].
> Â Â Â Â ^row
>
>
> with a XTFileReadStream, just rewrite this:
>
> MessageTally spyOn: [| stream separator |
> Â Â Â Â separator := $;.
> Â Â Â Â stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
> 'Data.csv') reading encoding: #utf8) "rest reading".
> Â Â Â Â (stream ending: Character cr) slicing collect: [:line |
> Â Â Â Â Â Â Â Â | row char nextValue noValue |
> Â Â Â Â Â Â Â Â row := OrderedCollection new.
> Â Â Â Â Â Â Â Â nextValue := (WideString new: 32) writing.
> Â Â Â Â Â Â Â Â noValue := true.
>
> Â Â Â Â Â Â Â Â [[[(char := line get) isSeparator] whileTrue.
> Â Â Â Â Â Â Â Â noValue := false.
> Â Â Â Â Â Â Â Â char = $" ifTrue:
> Â Â Â Â Â Â Â Â Â Â Â Â [[nextValue write: (line ending: $") rest.
> Â Â Â Â Â Â Â Â Â Â Â Â (char := line get) = $"] whileTrue:
> Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â [nextValue put: $"].
> Â Â Â Â Â Â Â Â Â Â Â Â [char = separator] whileFalse: [char := line get]].
> Â Â Â Â Â Â Â Â [char = separator]
> Â Â Â Â Â Â Â Â Â Â Â Â whileFalse:
> Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â [nextValue put: char.
> Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â char := line get].
> Â Â Â Â Â Â Â Â row add: nextValue conclusion.
> Â Â Â Â Â Â Â Â nextValue := (WideString new: 32) writing] repeat]
> Â Â Â Â Â Â Â Â Â Â Â Â on: Incomplete do: [:exc | ].
>
> Â Â Â Â Â Â Â Â noValue ifFalse: [row add: nextValue conclusion].
> Â Â Â Â Â Â Â Â row]]
>
> For some reason, this was super slow (>20s), though my previous
> attempt without slicing was around 3.0s...
Ah Ah, replacing contentsSpecies with ^WideString in
XTEncodeRead/WriteStream took it back to 3.2s.
Nicolas
> (so you noticed the commented "rest reading" whose purpose is to load
> the whole file in memory)
>
> The major advantage is that you can just replace #collect: with
> #collecting: and get a stream of rows. Lazy is cool.
>
> Nicolas
>
> 2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
>> 2011/11/17 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
>>> 2011/11/17 Dennis Schetinin <chaetal(a)gmail.com>:
>>>> So, the answer for the question "Is there a simple way to improve the
>>>> performance?" would be "No", right?
>>>>
>>>
>>> No, the bottle neck is CSVParser as Henrik and Levente already noticed.
>>>
>>> My machine executes your code in 21.7s
>>>
>>> So I decided to play with SqueaXTream http://www.squeaksource.com/XTream/
>>> and inlined row decoding in a single method:
>>>
>>> CSVParser>>nextRow
>>> Â Â Â Â | row char line nextValue |
>>> Â Â Â Â line := stream nextLine readXtream.
>>> Â Â Â Â row := OrderedCollection new.
>>> Â Â Â Â nextValue := (WideString new: 32) writeXtream.
>>> Â Â Â Â line endOfStreamAction: [^row add: (nextValue contents); yourself].
>>> Â Â Â Â [[(char := line next) isSeparator] whileTrue.
>>> Â Â Â Â char = $" ifTrue:
>>> Â Â Â Â Â Â Â Â [[nextValue nextPutAll: (stream upTo: $").
>>
>> Err, it's line upTo: $"
>>
>>> Â Â Â Â Â Â Â Â line peekFor: $"] whileTrue:
>>> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: $"].
>>> Â Â Â Â Â Â Â Â [(char := line next) = separator] whileFalse].
>>> Â Â Â Â [char = separator]
>>> Â Â Â Â Â Â Â Â whileFalse:
>>> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: char.
>>> Â Â Â Â Â Â Â Â Â Â Â Â char := line next].
>>> Â Â Â Â row add: nextValue contents.
>>> Â Â Â Â nextValue reset] repeat
>>>
>>> This reduces execution time to 2.4 seconds
>>>
>>> I also tried to use SqueaXTream for the file itself (if you don't mind
>>> the ugly invocation);
>>> But it does not gain much because you file is essentially made of Wide chars...
>>> It would help only if you'd have more ASCII fields...
>>>
>>> MessageTally spyOn: [| stream |
>>> Â Â Â Â stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
>>> 'Data.csv') readXtream binary buffered ~> UTF8Decoder) buffered.
>>> Â Â Â Â [(CSVParser on: stream)
>>> Â Â Â Â Â Â Â Â useDelimiter: $;;
>>> Â Â Â Â Â Â Â Â rows ] ensure: [stream close]].
>>>
>>> This reduces execution time to 1.9 seconds though
>>>
>>> So I wouldn't say there's not plenty of room for performance ;)
>>>
>>> Nicolas
>>>
>>>> 2011/11/16 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
>>>>>
>>>>> I think that there are several issues:
>>>>>
>>>>> 1) Character with codePoint < 256 are unique, but Wider are not.
>>>>> The characters are not stored in the WideString (which instead stores
>>>>> the code points), but any enumeration using at: rather than wordAt:
>>>>> will create an overhead. I presume that it would be better to have a
>>>>> #codePointAt: defined as #byteAt: in ByteString and #wordAt: in
>>>>> WideString, so we could still define generic methods in String.
>>>>> 2) A lot of ByteString methods are optimized with ad hoc optional
>>>>> primitives, but very few WideString methods are.
>>>>> 3) Some of these optimized operations work with 256 element masks and
>>>>> cannot be generalized to WideString that easily
>>>>> 4) Some Stream operations that copy one Collection to the stream
>>>>> buffer or vice versa are also not optimized for mixed byte/word
>>>>> collections
>>>>> 5) UTF8 is optimized for ASCII anyway... And Squeak/Pharo converters
>>>>> also are (because they can use more primitives in this case)
>>>>> 6) Pharo ripped #squeakToUtf8 and #utf8ToSqueak which were maybe
>>>>> questionable from quality POV, but were not replaced from efficiency
>>>>> POV.
>>>>>
>>>>> I would say that a new VM with immediate Character values might help,
>>>>> but that won't be enough.
>>>>>
>>>>> Nicolas
>>>>>
>>>>> 2011/11/16 Sven Van Caekenberghe <sven(a)beta9.be>:
>>>>> > Dennis,
>>>>> >
>>>>> > On 16 Nov 2011, at 19:30, Dennis Schetinin wrote:
>>>>> >
>>>>> >> Hi!
>>>>> >>
>>>>> >> Parsing a UTF-8 CSV file with Russian symbols (using CSVParser), I
>>>>> >> found it really slow. It takes about 5 seconds to read a 53Kb file (136
>>>>> >> lines) and several minutes to parse a 3.4Mb file (about 8K lines).
>>>>> >>
>>>>> >> Profiler shows that nearly all the time is eaten by WideString >>
>>>>> >> at:put: which is invoked from WriteStream>>nextPut:. If I'm not mistaken,
>>>>> >> this means that <primitive: 66> there does not work in this case⦠So, to be
>>>>> >> short: is there a simple way to improve the performance? TIA
>>>>> >>
>>>>> >> --
>>>>> >> Dennis Schetinin
>>>>> >
>>>>> > I am interested in this problem.
>>>>> >
>>>>> > Would it be possible to provide a test file (not that I can read
>>>>> > Russian) and your test code ?
>>>>> >
>>>>> > I would like to try this myself.
>>>>> >
>>>>> > Sven
>>>>> >
>>>>>
>>>>
>>>>
>>>>
>>>> --
>>>> Dennis Schetinin
>>>>
>>>
>>
>
Nov. 18, 2011
Re: [Pharo-project] MVP: quite interesting.
by Germán Arduino
Yes, this is the implementation of Dolphin, really a bit tricky to
understand at start (at least was to me) but after understand how it works,
it's a really very good approach imho.
Cheers.
2011/11/17 Stéphane Ducasse <stephane.ducasse(a)inria.fr>
>
> http://aviadezra.blogspot.com/2007/07/twisting-mvp-triad-say-hello-to-mvpc.…
> sounds also interesting.
>
>
> On Nov 17, 2011, at 10:49 PM, Stéphane Ducasse wrote:
>
> >
> http://www.mimuw.edu.pl/~sl/teaching/00_01/Delfin_EC/Overviews/ModelViewPre…
> >
> > Stef
>
>
>
--
============================================
Germán S. Arduino <gsa @ arsol.net> Twitter: garduino
Arduino Software http://www.arduinosoftware.com
PasswordsPro http://www.passwordspro.com
Promoter http://www.arsol.biz
============================================
Nov. 18, 2011
Re: [Pharo-project] WideString performance
by Nicolas Cellier
Also note that the more mainXtream Xtreams at
http://www.squeaksource.com/Xtreams/ does not perform that bad...
This ones gets a score of 3.0 seconds with a regular MultiByteFileStream
nextRow
| row char line nextValue noValue |
line := stream nextLine reading.
row := OrderedCollection new.
nextValue := (WideString new: 32) writing.
noValue := true.
[[[(char := line get) isSeparator] whileTrue.
noValue := false.
char = $" ifTrue:
[[nextValue write: (line ending: $") rest.
(char := line get) = $"] whileTrue:
[nextValue put: $"].
[char = separator] whileFalse: [char := line get]].
[char = separator]
whileFalse:
[nextValue put: char.
char := line get].
row add: nextValue conclusion.
nextValue := (WideString new: 32) writing] repeat]
on: Incomplete do: [:exc | ].
noValue ifFalse: [row add: nextValue conclusion].
^row
with a XTFileReadStream, just rewrite this:
MessageTally spyOn: [| stream separator |
separator := $;.
stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
'Data.csv') reading encoding: #utf8) "rest reading".
(stream ending: Character cr) slicing collect: [:line |
| row char nextValue noValue |
row := OrderedCollection new.
nextValue := (WideString new: 32) writing.
noValue := true.
[[[(char := line get) isSeparator] whileTrue.
noValue := false.
char = $" ifTrue:
[[nextValue write: (line ending: $") rest.
(char := line get) = $"] whileTrue:
[nextValue put: $"].
[char = separator] whileFalse: [char := line get]].
[char = separator]
whileFalse:
[nextValue put: char.
char := line get].
row add: nextValue conclusion.
nextValue := (WideString new: 32) writing] repeat]
on: Incomplete do: [:exc | ].
noValue ifFalse: [row add: nextValue conclusion].
row]]
For some reason, this was super slow (>20s), though my previous
attempt without slicing was around 3.0s...
(so you noticed the commented "rest reading" whose purpose is to load
the whole file in memory)
The major advantage is that you can just replace #collect: with
#collecting: and get a stream of rows. Lazy is cool.
Nicolas
2011/11/18 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> 2011/11/17 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
>> 2011/11/17 Dennis Schetinin <chaetal(a)gmail.com>:
>>> So, the answer for the question "Is there a simple way to improve the
>>> performance?" would be "No", right?
>>>
>>
>> No, the bottle neck is CSVParser as Henrik and Levente already noticed.
>>
>> My machine executes your code in 21.7s
>>
>> So I decided to play with SqueaXTream http://www.squeaksource.com/XTream/
>> and inlined row decoding in a single method:
>>
>> CSVParser>>nextRow
>> Â Â Â Â | row char line nextValue |
>> Â Â Â Â line := stream nextLine readXtream.
>> Â Â Â Â row := OrderedCollection new.
>> Â Â Â Â nextValue := (WideString new: 32) writeXtream.
>> Â Â Â Â line endOfStreamAction: [^row add: (nextValue contents); yourself].
>> Â Â Â Â [[(char := line next) isSeparator] whileTrue.
>> Â Â Â Â char = $" ifTrue:
>> Â Â Â Â Â Â Â Â [[nextValue nextPutAll: (stream upTo: $").
>
> Err, it's line upTo: $"
>
>> Â Â Â Â Â Â Â Â line peekFor: $"] whileTrue:
>> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: $"].
>> Â Â Â Â Â Â Â Â [(char := line next) = separator] whileFalse].
>> Â Â Â Â [char = separator]
>> Â Â Â Â Â Â Â Â whileFalse:
>> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: char.
>> Â Â Â Â Â Â Â Â Â Â Â Â char := line next].
>> Â Â Â Â row add: nextValue contents.
>> Â Â Â Â nextValue reset] repeat
>>
>> This reduces execution time to 2.4 seconds
>>
>> I also tried to use SqueaXTream for the file itself (if you don't mind
>> the ugly invocation);
>> But it does not gain much because you file is essentially made of Wide chars...
>> It would help only if you'd have more ASCII fields...
>>
>> MessageTally spyOn: [| stream |
>> Â Â Â Â stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
>> 'Data.csv') readXtream binary buffered ~> UTF8Decoder) buffered.
>> Â Â Â Â [(CSVParser on: stream)
>> Â Â Â Â Â Â Â Â useDelimiter: $;;
>> Â Â Â Â Â Â Â Â rows ] ensure: [stream close]].
>>
>> This reduces execution time to 1.9 seconds though
>>
>> So I wouldn't say there's not plenty of room for performance ;)
>>
>> Nicolas
>>
>>> 2011/11/16 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
>>>>
>>>> I think that there are several issues:
>>>>
>>>> 1) Character with codePoint < 256 are unique, but Wider are not.
>>>> The characters are not stored in the WideString (which instead stores
>>>> the code points), but any enumeration using at: rather than wordAt:
>>>> will create an overhead. I presume that it would be better to have a
>>>> #codePointAt: defined as #byteAt: in ByteString and #wordAt: in
>>>> WideString, so we could still define generic methods in String.
>>>> 2) A lot of ByteString methods are optimized with ad hoc optional
>>>> primitives, but very few WideString methods are.
>>>> 3) Some of these optimized operations work with 256 element masks and
>>>> cannot be generalized to WideString that easily
>>>> 4) Some Stream operations that copy one Collection to the stream
>>>> buffer or vice versa are also not optimized for mixed byte/word
>>>> collections
>>>> 5) UTF8 is optimized for ASCII anyway... And Squeak/Pharo converters
>>>> also are (because they can use more primitives in this case)
>>>> 6) Pharo ripped #squeakToUtf8 and #utf8ToSqueak which were maybe
>>>> questionable from quality POV, but were not replaced from efficiency
>>>> POV.
>>>>
>>>> I would say that a new VM with immediate Character values might help,
>>>> but that won't be enough.
>>>>
>>>> Nicolas
>>>>
>>>> 2011/11/16 Sven Van Caekenberghe <sven(a)beta9.be>:
>>>> > Dennis,
>>>> >
>>>> > On 16 Nov 2011, at 19:30, Dennis Schetinin wrote:
>>>> >
>>>> >> Hi!
>>>> >>
>>>> >> Parsing a UTF-8 CSV file with Russian symbols (using CSVParser), I
>>>> >> found it really slow. It takes about 5 seconds to read a 53Kb file (136
>>>> >> lines) and several minutes to parse a 3.4Mb file (about 8K lines).
>>>> >>
>>>> >> Profiler shows that nearly all the time is eaten by WideString >>
>>>> >> at:put: which is invoked from WriteStream>>nextPut:. If I'm not mistaken,
>>>> >> this means that <primitive: 66> there does not work in this case⦠So, to be
>>>> >> short: is there a simple way to improve the performance? TIA
>>>> >>
>>>> >> --
>>>> >> Dennis Schetinin
>>>> >
>>>> > I am interested in this problem.
>>>> >
>>>> > Would it be possible to provide a test file (not that I can read
>>>> > Russian) and your test code ?
>>>> >
>>>> > I would like to try this myself.
>>>> >
>>>> > Sven
>>>> >
>>>>
>>>
>>>
>>>
>>> --
>>> Dennis Schetinin
>>>
>>
>
Nov. 17, 2011
Re: [Pharo-project] WideString performance
by Nicolas Cellier
2011/11/17 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>:
> 2011/11/17 Dennis Schetinin <chaetal(a)gmail.com>:
>> So, the answer for the question "Is there a simple way to improve the
>> performance?" would be "No", right?
>>
>
> No, the bottle neck is CSVParser as Henrik and Levente already noticed.
>
> My machine executes your code in 21.7s
>
> So I decided to play with SqueaXTream http://www.squeaksource.com/XTream/
> and inlined row decoding in a single method:
>
> CSVParser>>nextRow
> Â Â Â Â | row char line nextValue |
> Â Â Â Â line := stream nextLine readXtream.
> Â Â Â Â row := OrderedCollection new.
> Â Â Â Â nextValue := (WideString new: 32) writeXtream.
> Â Â Â Â line endOfStreamAction: [^row add: (nextValue contents); yourself].
> Â Â Â Â [[(char := line next) isSeparator] whileTrue.
> Â Â Â Â char = $" ifTrue:
> Â Â Â Â Â Â Â Â [[nextValue nextPutAll: (stream upTo: $").
Err, it's line upTo: $"
> Â Â Â Â Â Â Â Â line peekFor: $"] whileTrue:
> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: $"].
> Â Â Â Â Â Â Â Â [(char := line next) = separator] whileFalse].
> Â Â Â Â [char = separator]
> Â Â Â Â Â Â Â Â whileFalse:
> Â Â Â Â Â Â Â Â Â Â Â Â [nextValue nextPut: char.
> Â Â Â Â Â Â Â Â Â Â Â Â char := line next].
> Â Â Â Â row add: nextValue contents.
> Â Â Â Â nextValue reset] repeat
>
> This reduces execution time to 2.4 seconds
>
> I also tried to use SqueaXTream for the file itself (if you don't mind
> the ugly invocation);
> But it does not gain much because you file is essentially made of Wide chars...
> It would help only if you'd have more ASCII fields...
>
> MessageTally spyOn: [| stream |
> Â Â Â Â stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
> 'Data.csv') readXtream binary buffered ~> UTF8Decoder) buffered.
> Â Â Â Â [(CSVParser on: stream)
> Â Â Â Â Â Â Â Â useDelimiter: $;;
> Â Â Â Â Â Â Â Â rows ] ensure: [stream close]].
>
> This reduces execution time to 1.9 seconds though
>
> So I wouldn't say there's not plenty of room for performance ;)
>
> Nicolas
>
>> 2011/11/16 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
>>>
>>> I think that there are several issues:
>>>
>>> 1) Character with codePoint < 256 are unique, but Wider are not.
>>> The characters are not stored in the WideString (which instead stores
>>> the code points), but any enumeration using at: rather than wordAt:
>>> will create an overhead. I presume that it would be better to have a
>>> #codePointAt: defined as #byteAt: in ByteString and #wordAt: in
>>> WideString, so we could still define generic methods in String.
>>> 2) A lot of ByteString methods are optimized with ad hoc optional
>>> primitives, but very few WideString methods are.
>>> 3) Some of these optimized operations work with 256 element masks and
>>> cannot be generalized to WideString that easily
>>> 4) Some Stream operations that copy one Collection to the stream
>>> buffer or vice versa are also not optimized for mixed byte/word
>>> collections
>>> 5) UTF8 is optimized for ASCII anyway... And Squeak/Pharo converters
>>> also are (because they can use more primitives in this case)
>>> 6) Pharo ripped #squeakToUtf8 and #utf8ToSqueak which were maybe
>>> questionable from quality POV, but were not replaced from efficiency
>>> POV.
>>>
>>> I would say that a new VM with immediate Character values might help,
>>> but that won't be enough.
>>>
>>> Nicolas
>>>
>>> 2011/11/16 Sven Van Caekenberghe <sven(a)beta9.be>:
>>> > Dennis,
>>> >
>>> > On 16 Nov 2011, at 19:30, Dennis Schetinin wrote:
>>> >
>>> >> Hi!
>>> >>
>>> >> Parsing a UTF-8 CSV file with Russian symbols (using CSVParser), I
>>> >> found it really slow. It takes about 5 seconds to read a 53Kb file (136
>>> >> lines) and several minutes to parse a 3.4Mb file (about 8K lines).
>>> >>
>>> >> Profiler shows that nearly all the time is eaten by WideString >>
>>> >> at:put: which is invoked from WriteStream>>nextPut:. If I'm not mistaken,
>>> >> this means that <primitive: 66> there does not work in this case⦠So, to be
>>> >> short: is there a simple way to improve the performance? TIA
>>> >>
>>> >> --
>>> >> Dennis Schetinin
>>> >
>>> > I am interested in this problem.
>>> >
>>> > Would it be possible to provide a test file (not that I can read
>>> > Russian) and your test code ?
>>> >
>>> > I would like to try this myself.
>>> >
>>> > Sven
>>> >
>>>
>>
>>
>>
>> --
>> Dennis Schetinin
>>
>
Nov. 17, 2011
Re: [Pharo-project] OpenGL in Pharo?
by Pavel Krivanek
I would like to ask on the future of FFI technologies in Pharo. What will be the prefered project? Alien, NativeBoost or both? That means what projects will have support in the "official" VM for all main platforms?
Some week ago I played with FFI calls on Linux and the only project that worked well was NativeBoost. The Alien failed to locate the libraries (it was 32bit VM on 64bit system, possible source of complacations).
Cheers
-- Pavel
17.11.2011 v 18:51, Stéphane Ducasse <stephane.ducasse(a)inria.fr>:
> Hi michele
>
> Do not expect something like Jun because this is 10 or 15 years. Now I saw today a cairo graphics displayed on the machine of Javier :).
> This is an example for the Alien chapter of the new pharo book (quite draft).
>
> I hope that I will get somebody will work on a minimal code city for moose based on a opengl solution.
> Now if people have more cycles it will go faster. I hope igor will release soon the opengl binding for mac.
>
> Stef
> <alien.pdf>
>
>
> On Nov 17, 2011, at 8:52 AM, Michele Lanza wrote:
>
>> Dear all,
>> I wanted to ask whether there are some news I missed with respect to having a decent OpenGL framework in Pharo. I checked the following page:
>>
>> http://book.pharo-project.org/book/LanguageAndLibraries/3DGraphicsAndOpenGL/
>>
>> ..but it sounds all a bit outdated and not something I can rely on in the long run. Ideally there should be something like the Jun framework (available for VisualWorks), but when I met the Jun maker at ICSE and I asked him whether there were any plans of porting Jun to Pharo he told me that it was not in their plans. Long talk, short sense: Is anyone working on OpenGL for Pharo?
>>
>> Cheers
>>
>> Michele
>
Nov. 17, 2011
Re: [Pharo-project] WideString performance
by Nicolas Cellier
2011/11/17 Dennis Schetinin <chaetal(a)gmail.com>:
> So, the answer for the question "Is there a simple way to improve the
> performance?" would be "No", right?
>
No, the bottle neck is CSVParser as Henrik and Levente already noticed.
My machine executes your code in 21.7s
So I decided to play with SqueaXTream http://www.squeaksource.com/XTream/
and inlined row decoding in a single method:
CSVParser>>nextRow
| row char line nextValue |
line := stream nextLine readXtream.
row := OrderedCollection new.
nextValue := (WideString new: 32) writeXtream.
line endOfStreamAction: [^row add: (nextValue contents); yourself].
[[(char := line next) isSeparator] whileTrue.
char = $" ifTrue:
[[nextValue nextPutAll: (stream upTo: $").
line peekFor: $"] whileTrue:
[nextValue nextPut: $"].
[(char := line next) = separator] whileFalse].
[char = separator]
whileFalse:
[nextValue nextPut: char.
char := line next].
row add: nextValue contents.
nextValue reset] repeat
This reduces execution time to 2.4 seconds
I also tried to use SqueaXTream for the file itself (if you don't mind
the ugly invocation);
But it does not gain much because you file is essentially made of Wide chars...
It would help only if you'd have more ASCII fields...
MessageTally spyOn: [| stream |
stream := (((FileDirectory on: '/Users/nicolas/Downloads/') /
'Data.csv') readXtream binary buffered ~> UTF8Decoder) buffered.
[(CSVParser on: stream)
useDelimiter: $;;
rows ] ensure: [stream close]].
This reduces execution time to 1.9 seconds though
So I wouldn't say there's not plenty of room for performance ;)
Nicolas
> 2011/11/16 Nicolas Cellier <nicolas.cellier.aka.nice(a)gmail.com>
>>
>> I think that there are several issues:
>>
>> 1) Character with codePoint < 256 are unique, but Wider are not.
>> The characters are not stored in the WideString (which instead stores
>> the code points), but any enumeration using at: rather than wordAt:
>> will create an overhead. I presume that it would be better to have a
>> #codePointAt: defined as #byteAt: in ByteString and #wordAt: in
>> WideString, so we could still define generic methods in String.
>> 2) A lot of ByteString methods are optimized with ad hoc optional
>> primitives, but very few WideString methods are.
>> 3) Some of these optimized operations work with 256 element masks and
>> cannot be generalized to WideString that easily
>> 4) Some Stream operations that copy one Collection to the stream
>> buffer or vice versa are also not optimized for mixed byte/word
>> collections
>> 5) UTF8 is optimized for ASCII anyway... And Squeak/Pharo converters
>> also are (because they can use more primitives in this case)
>> 6) Pharo ripped #squeakToUtf8 and #utf8ToSqueak which were maybe
>> questionable from quality POV, but were not replaced from efficiency
>> POV.
>>
>> I would say that a new VM with immediate Character values might help,
>> but that won't be enough.
>>
>> Nicolas
>>
>> 2011/11/16 Sven Van Caekenberghe <sven(a)beta9.be>:
>> > Dennis,
>> >
>> > On 16 Nov 2011, at 19:30, Dennis Schetinin wrote:
>> >
>> >> Hi!
>> >>
>> >> Parsing a UTF-8 CSV file with Russian symbols (using CSVParser), I
>> >> found it really slow. It takes about 5 seconds to read a 53Kb file (136
>> >> lines) and several minutes to parse a 3.4Mb file (about 8K lines).
>> >>
>> >> Profiler shows that nearly all the time is eaten by WideString >>
>> >> at:put: which is invoked from WriteStream>>nextPut:. If I'm not mistaken,
>> >> this means that <primitive: 66> there does not work in this case⦠So, to be
>> >> short: is there a simple way to improve the performance? TIA
>> >>
>> >> --
>> >> Dennis Schetinin
>> >
>> > I am interested in this problem.
>> >
>> > Would it be possible to provide a test file (not that I can read
>> > Russian) and your test code ?
>> >
>> > I would like to try this myself.
>> >
>> > Sven
>> >
>>
>
>
>
> --
> Dennis Schetinin
>
Nov. 17, 2011
Re: [Pharo-project] how to get a string read stream from filesystem
by Norbert Hartl
Am 17.11.2011 um 22:01 schrieb Tudor Girba:
> Hi,
>
> I am using Filesystem, and I would like to get a read stream that provides a string instead of bytes.
>
> If I do "aReference readStream", I get a FSReadStream which is a byte stream.
>
> To get to the string, I now do the ugly: "aReference readStream contents asString"
>
> Is there a cleverer way?
>
I don't know FileSystem. But in any case a bytestream is correct. In order to get a string you need to decode it. Only for ascii this is null codec. For any other encoded file you won't get what you want. So, using a textconverter with file encoding might get you what you want.
Norbert
Nov. 17, 2011