Pharo-dev
By thread
pharo-dev@lists.pharo.org
By month
Messages by month
- ----- 2026 -----
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2008 -----
- December
- November
- October
- September
- August
- July
- June
- May
- 1 participants
- 144621 messages
WhatsUp from: 2014-03-03 until: 2014-03-16
by seaside@rmod.lille.inria.fr
Hi! We're sending this automatic email twice a month, to give the community an opportunity to easily know what's happening and to coordinate efforts. Just answer informally, and feel free to spawn discussions thereafter!
### Here's what I've been up to since the last WhatsUp:
- $HEROIC_ACHIEVEMENTS_OR_DISMAL_FAILURES_OR_SIMPLE_BORING_NECESSARY_TASKS
### What's next, until 2014-03-16 (*):
- $NEXT_STEPS_TOWARDS_WORLD_DOMINATION
(*) we'll be expecting results by then ;)
March 3, 2014
Re: [Pharo-dev] [Pharo-business] [ANN] DBPedia: Query Wikipedia from Pharo
by Hernán Morales Durand
2014-03-02 21:22 GMT-03:00 Alexandre Bergel <alexandre.bergel(a)me.com>:
> Iâve just tried and it works pretty well! Impressive!
>
> Below I describe a small example that fetches some data about the US
> Universities from DBPedia and visualize them using Roassal2.
>
> Pick a fresh 3.0 image.
>
> First, you need to load Hernán work, Svenâs NeoJSON, and Roassal 2 (If you
> are using a Moose Image, there is no need to load Roassal2 since it is
> already in):
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
> Gofer it
> smalltalkhubUser: 'SvenVanCaekenberghe' project: 'Neo';
> package: 'ConfigurationOfNeoJSON';
> load.
> ((Smalltalk at: #ConfigurationOfNeoJSON) load).
>
> Gofer it
> smalltalkhubUser: 'hernan' project: 'DBPedia';
> package: 'DBPedia';
> load.
>
> Gofer it
> smalltalkhubUser: 'ObjectProfile' project: 'Roassal2';
> package: 'ConfigurationOfRoassal2';
> load.
> ((Smalltalk at: #ConfigurationOfRoassal2) loadBleedingEdge).
>
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
>
> Using Roassal2, I was able to render some data extracted from dbpedia:
>
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
> | map locations rawData rawData2 rawData3 |
> map := RTMapBuilder new.
>
> map countries: #('UnitedStates' 'Canada' 'Mexico').
> map color: Color veryVeryLightGray.
>
> rawData := DBPediaSearch universitiesInUS.
> rawData2 := ((NeoJSONReader fromString: rawData) at: #results) at:
> #bindings.
> rawData3 := rawData2 select: [ :d | d keys includesAll: #('label' 'long'
> 'lat') ] thenCollect: [ :d | { (Float readFrom: ((d at: 'long') at:
> 'value')) . (Float readFrom: ((d at: 'lat') at: 'value')) . (d at: 'label'
> ) at: 'value' } ].
>
>
> locations := rawData3.
> locations do: [ :array |
> map cities addCityNamed: array third location: array second @ array first
> ].
> map cities shape size: 8; color: (Color blue alpha: 0.03).
> map cities: (locations collect: #third).
>
> map scale: 2.
>
> map render.
> map view openInWindowSized: 1000 @ 500.
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
>
> This is what you get:
>
>
> This is a small example. Naturally, adding popup for locations is trivial
> to add.
>
> I have described this on our Facebook page:
>
> https://www.facebook.com/ObjectProfile/photos/a.341189379300999.82969.34054…
>
>
Super cool!! Thanks for sharing the nice mapping.
Hernán, since SPARQL is a bit obscure,
>
Absolutely, SPARQL is like the Assembler of the web.
> it would be great if you could add some more example, and also, how to
> parametrize the examples. For example, now we can get data for the US, how
> to modify your example to get them for France or Chile?
>
>
Ok, uploaded an updated version. I have parametrized the query triplets as
#universitiesIn: englishCountryName,
DBPediaSearch universitiesIn: 'France'.
DBPediaSearch universitiesIn: 'Chile'.
I have to admit I am still learning SPARQL, but the more queries I execute,
the more I discover linked data structure, so I will add convenience
methods for easy parsing results.
Cheers,
Hernán
March 3, 2014
Re: [Pharo-dev] [Pharo-business] [ANN] DBPedia: Query Wikipedia from Pharo
by Alexandre Bergel
A followup from the previous post. Fetching country population and charting them using GraphET:
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
| query data diagram |
query := DBPediaSearch new
setJsonFormat;
setDebugOn;
timeout: 3000;
query: 'SELECT DISTINCT ?name ?population
WHERE {
?country a dbpedia-owl:Country .
?country rdfs:label ?name .
FILTER
(langMatches(lang(?name), "en"))
values ?hasPopulation { dbpprop:populationEstimatedbpprop:populationCensus }
OPTIONAL { ?country ?hasPopulation ?population }
FILTER (isNumeric(?population))
FILTER NOT EXISTS { ?country dbpedia-owl:dissolutionYear ?yearEnd } { ?country dbpprop:iso3166code ?code . }
UNION { ?country dbpprop:iso31661Alpha ?code . }
UNION { ?country dbpprop:countryCode ?code . }
UNION { ?country a yago:MemberStatesOfTheUnitedNations . }}';
execute.
data := (((NeoJSONReader fromString: query) at:#results) at: #bindings) collect: [ :entry | Array with: ( (entry at: #name) at: #value ) with: ( (entry at: #population) at: #value ) asInteger ].
"Use GraphET to render all this"
diagram := GETDiagramBuilder new.
diagram verticalBarDiagram
models: (data reverseSortedAs: #second);
y: #second;
regularAxisAsInteger;
titleLabel: 'Size of countries';
yAxisLabel: 'Population'.
diagram interaction popupText.
diagram open
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
Cheers,
Alexandre
On Mar 2, 2014, at 9:22 PM, Alexandre Bergel <alexandre.bergel(a)me.com> wrote:
> Iâve just tried and it works pretty well! Impressive!
>
> Below I describe a small example that fetches some data about the US Universities from DBPedia and visualize them using Roassal2.
>
> Pick a fresh 3.0 image.
>
> First, you need to load Hernán work, Svenâs NeoJSON, and Roassal 2 (If you are using a Moose Image, there is no need to load Roassal2 since it is already in):
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
> Gofer it
> smalltalkhubUser: 'SvenVanCaekenberghe' project: 'Neo';
> package: 'ConfigurationOfNeoJSON';
> load.
> ((Smalltalk at: #ConfigurationOfNeoJSON) load).
>
> Gofer it
> smalltalkhubUser: 'hernan' project: 'DBPedia';
> package: 'DBPedia';
> load.
>
> Gofer it
> smalltalkhubUser: 'ObjectProfile' project: 'Roassal2';
> package: 'ConfigurationOfRoassal2';
> load.
> ((Smalltalk at: #ConfigurationOfRoassal2) loadBleedingEdge).
>
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
>
> Using Roassal2, I was able to render some data extracted from dbpedia:
>
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
> | map locations rawData rawData2 rawData3 |
> map := RTMapBuilder new.
>
> map countries: #('UnitedStates' 'Canada' 'Mexico').
> map color: Color veryVeryLightGray.
>
> rawData := DBPediaSearch universitiesInUS.
> rawData2 := ((NeoJSONReader fromString: rawData) at: #results) at: #bindings.
> rawData3 := rawData2 select: [ :d | d keys includesAll: #('label' 'long' 'lat') ] thenCollect: [ :d | { (Float readFrom: ((d at: 'long') at: 'value')) . (Float readFrom: ((d at: 'lat') at: 'value')) . (d at: 'label' ) at: 'value' } ].
>
>
> locations := rawData3.
> locations do: [ :array |
> map cities addCityNamed: array third location: array second @ array first ].
> map cities shape size: 8; color: (Color blue alpha: 0.03).
> map cities: (locations collect: #third).
>
> map scale: 2.
>
> map render.
> map view openInWindowSized: 1000 @ 500.
> -=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
>
> This is what you get:
>
> <Screen Shot 2014-03-02 at 9.09.57 PM.png>
>
> This is a small example. Naturally, adding popup for locations is trivial to add.
>
> I have described this on our Facebook page:
> https://www.facebook.com/ObjectProfile/photos/a.341189379300999.82969.34054…
>
> Hernán, since SPARQL is a bit obscure, it would be great if you could add some more example, and also, how to parametrize the examples. For example, now we can get data for the US, how to modify your example to get them for France or Chile?
>
> Cheers,
> Alexandre
>
>
> On Mar 2, 2014, at 3:43 PM, Hernán Morales Durand <hernan.morales(a)gmail.com> wrote:
>
>> I have uploaded a new configuration so you can query the english Wikipedia dataset from Pharo 3 using SPARQL. Some examples follow:
>>
>> 1) Retrieve in JSON movies from the beautiful Julianne Moore:
>>
>> | jsonResults |
>> jsonResults := DBPediaSearch new
>> setJsonFormat;
>> timeout: 5000;
>> query: 'SELECT DISTINCT ?filmName WHERE {
>> ?film foaf:name ?filmName .
>> ?film dbpedia-owl:starring ?actress .
>> ?actress foaf:name ?name.
>> FILTER(contains(?name, "Julianne"))
>> FILTER(contains(?name, "Moore"))
>> }';
>> execute
>>
>> To actually get only the titles using NeoJSON:
>>
>> ((((NeoJSONReader fromString: jsonResults) at: #results) at: #bindings)
>> collect: [ : entry | entry at: #filmName ]) collect: [ : movie | movie at: #value ]
>>
>>
>> 2) Retrieve in XML which genre plays those crazy Dream Theater guys :
>>
>> DBPediaSearch new
>> setXmlFormat;
>> setDebugOn;
>> timeout: 5000;
>> query: 'SELECT DISTINCT ?genreLabel
>> WHERE {
>> ?resource dbpprop:genre ?genre.
>> ?resource rdfs:label "Dream Theater"@en.
>> ?genre rdfs:label ?genreLabel
>> FILTER (lang(?genreLabel)="en")
>> }
>> LIMIT 100';
>> execute
>>
>> More examples are available in DBPediaSearch class side. You can install it from the Configuration Browser.
>> If you want to contribute, just ask me and you will be added as contributor.
>> Best regards,
>>
>> Hernán
>>
>> _______________________________________________
>> Pharo-business mailing list
>> Pharo-business(a)lists.pharo.org
>> http://lists.pharo.org/mailman/listinfo/pharo-business_lists.pharo.org
>
> --
> _,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:
> Alexandre Bergel http://www.bergel.eu
> ^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;.
>
>
>
--
_,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:
Alexandre Bergel http://www.bergel.eu
^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;.
March 3, 2014
Re: [Pharo-dev] [Pharo-business] [ANN] DBPedia: Query Wikipedia from Pharo
by Alexandre Bergel
Iâve just tried and it works pretty well! Impressive!
Below I describe a small example that fetches some data about the US Universities from DBPedia and visualize them using Roassal2.
Pick a fresh 3.0 image.
First, you need to load Hernán work, Svenâs NeoJSON, and Roassal 2 (If you are using a Moose Image, there is no need to load Roassal2 since it is already in):
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
Gofer it
smalltalkhubUser: 'SvenVanCaekenberghe' project: 'Neo';
package: 'ConfigurationOfNeoJSON';
load.
((Smalltalk at: #ConfigurationOfNeoJSON) load).
Gofer it
smalltalkhubUser: 'hernan' project: 'DBPedia';
package: 'DBPedia';
load.
Gofer it
smalltalkhubUser: 'ObjectProfile' project: 'Roassal2';
package: 'ConfigurationOfRoassal2';
load.
((Smalltalk at: #ConfigurationOfRoassal2) loadBleedingEdge).
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
Using Roassal2, I was able to render some data extracted from dbpedia:
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
| map locations rawData rawData2 rawData3 |
map := RTMapBuilder new.
map countries: #('UnitedStates' 'Canada' 'Mexico').
map color: Color veryVeryLightGray.
rawData := DBPediaSearch universitiesInUS.
rawData2 := ((NeoJSONReader fromString: rawData) at: #results) at: #bindings.
rawData3 := rawData2 select: [ :d | d keys includesAll: #('label' 'long' 'lat') ] thenCollect: [ :d | { (Float readFrom: ((d at: 'long') at: 'value')) . (Float readFrom: ((d at: 'lat') at: 'value')) . (d at: 'label' ) at: 'value' } ].
locations := rawData3.
locations do: [ :array |
map cities addCityNamed: array third location: array second @ array first ].
map cities shape size: 8; color: (Color blue alpha: 0.03).
map cities: (locations collect: #third).
map scale: 2.
map render.
map view openInWindowSized: 1000 @ 500.
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
This is what you get:
This is a small example. Naturally, adding popup for locations is trivial to add.
I have described this on our Facebook page:
https://www.facebook.com/ObjectProfile/photos/a.341189379300999.82969.34054…
Hernán, since SPARQL is a bit obscure, it would be great if you could add some more example, and also, how to parametrize the examples. For example, now we can get data for the US, how to modify your example to get them for France or Chile?
Cheers,
Alexandre
On Mar 2, 2014, at 3:43 PM, Hernán Morales Durand <hernan.morales(a)gmail.com> wrote:
> I have uploaded a new configuration so you can query the english Wikipedia dataset from Pharo 3 using SPARQL. Some examples follow:
>
> 1) Retrieve in JSON movies from the beautiful Julianne Moore:
>
> | jsonResults |
> jsonResults := DBPediaSearch new
> setJsonFormat;
> timeout: 5000;
> query: 'SELECT DISTINCT ?filmName WHERE {
> ?film foaf:name ?filmName .
> ?film dbpedia-owl:starring ?actress .
> ?actress foaf:name ?name.
> FILTER(contains(?name, "Julianne"))
> FILTER(contains(?name, "Moore"))
> }';
> execute
>
> To actually get only the titles using NeoJSON:
>
> ((((NeoJSONReader fromString: jsonResults) at: #results) at: #bindings)
> collect: [ : entry | entry at: #filmName ]) collect: [ : movie | movie at: #value ]
>
>
> 2) Retrieve in XML which genre plays those crazy Dream Theater guys :
>
> DBPediaSearch new
> setXmlFormat;
> setDebugOn;
> timeout: 5000;
> query: 'SELECT DISTINCT ?genreLabel
> WHERE {
> ?resource dbpprop:genre ?genre.
> ?resource rdfs:label "Dream Theater"@en.
> ?genre rdfs:label ?genreLabel
> FILTER (lang(?genreLabel)="en")
> }
> LIMIT 100';
> execute
>
> More examples are available in DBPediaSearch class side. You can install it from the Configuration Browser.
> If you want to contribute, just ask me and you will be added as contributor.
> Best regards,
>
> Hernán
>
> _______________________________________________
> Pharo-business mailing list
> Pharo-business(a)lists.pharo.org
> http://lists.pharo.org/mailman/listinfo/pharo-business_lists.pharo.org
--
_,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:
Alexandre Bergel http://www.bergel.eu
^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;._,.;:~^~:;.
March 3, 2014
Re: [Pharo-dev] understanding postgresv2
by Yanni Chiu
On 02/03/2014 11:34 AM, Tudor Girba wrote:
>
> - Why is result an instance variable in PGConnection? Making it a
> variable always returns the same object when executing a query and that
> is a bit of a pain.
Because the PGConnection is designed as an active object, controlled by
a state machine. Depending on the connection state, the result set
associated with the connection is either valid or not valid.
The layer above PGConnection should not hold onto any PGResult directly,
since it's internal data structure of the Postgres client. The object
mapping layer would be expected to copy data values from the result rows.
> - Why does the PGResult have the possibility of holding multiple
> PGResultSets? When is it possible to have multiple at the same time?
> (when you execute a query, the result is being initialized)
IIRC, something like:
connection execute 'select * from foo1; select * from foo2'
will get you multiple result sets.
> - When running something like
> connection execute: 'select * from ...'
> the PGResultSet already includes all rows of the query. Is it not
> possible to have a stream-like functionality in which the actual rows
> are retrieved only on demand? (a similar functionality exists in DBXTalk)
IIUC (it's been a long time since I looked at it), the V3 protocol has
this streaming support. In the V2 protocol, I think you have to read and
discard the results. As for DBXTalk, it's likely using the V3 protocol.
> - What is the difference between PGAsciiRow and PGDataRow?
The message format is documented at:
http://www.postgresql.org/docs/7.1/static/protocol-message-formats.html
I see an AsciiRow and a BinaryRow there, but no DataRow. So, I don't
know where DataRow comes from, unless it's from an even older version of
the protocol docs (original development was done on Postgres version
6.4, IIRC).
March 2, 2014
Re: [Pharo-dev] hashMultiply
by Andres Valloud
Once the code is loaded, from the Tools menu use Hash Analysis Tool.
There's a manual below, and also the Fundamentals book has a somewhat in
depth discussion on how it works.
ftp://sqrmax.us.to/pub/Smalltalk/Papers/Hash%20Analysis%20Tool.pdf
On 3/2/14 15:06 , phil(a)highoctane.be wrote:
> I have found the tools in the Cincom public repo.
>
> Now, I have to see how this works in VW, I am not that proficient with it.
>
> Phil
>
>
>
> On Wed, Feb 26, 2014 at 8:55 AM, phil(a)highoctane.be
> <mailto:phil@highoctane.be> <phil(a)highoctane.be
> <mailto:phil@highoctane.be>> wrote:
>
> Andres,
>
> Thanks for the insights.
>
> hash quality is indeed an key factor. At least, my mind is somewhat
> grasping this hashing field a bit better.
>
> I'll have a look at the tools, I haven't used them yet.
>
> And a shot at the ASM version with NativeBoost in Pharo.
>
> For what occurs in modern CPUs, well, no. Surprising to see that a
> mul would be faster than a shr or shl. How comes?
> I used to be ok with these things when I was writing demoscene code
> a loong time ago but I'd need an extremely serious refresh.
>
> As a side note, there is a huge uptake in the
> BigData/mapreduce/hadoop environment where Smalltalk is sorely
> absent. Scala seems to fill the void on the JVM.
> There hashing is quite key, to remap all of the mapping phase
> results to the reduce nodes. I am surprised to see that Smalltalk
> vendors haven't jumped in that space.
>
> Phil
>
>
> On Wed, Feb 26, 2014 at 2:04 AM, Andres Valloud
> <avalloud(a)smalltalk.comcastbiz.net
> <mailto:avalloud@smalltalk.comcastbiz.net>> wrote:
>
> Hello...
>
> On 2/25/14 1:17 , phil(a)highoctane.be <mailto:phil@highoctane.be>
> wrote:
>
> I am currently reading through the Hashing in Smalltalk book
> (http://www.lulu.com/shop/__andres-valloud/hashing-in-__smalltalk-theory-and…
> <http://www.lulu.com/shop/andres-valloud/hashing-in-smalltalk-theory-and-pra…>__)
> and, my head hurting notwithstanding, there are indeed a ton
> of gems in
> this system. As he mentions, doing the exercises brings a
> lot of extra :-)
>
>
> :) thank you.
>
> When going to 64-bit, and with the new ObjectMemory scheme,
> I guess a
> couple of identity hashing functions will come under scrutiny.
>
> e.g.
>
> SmallInteger>>hashMultiply
> | low |
>
> low := self bitAnd: 16383.
> ^(16r260D * low + ((16r260D * (self bitShift: -14) +
> (16r0065 * low)
> bitAnd: 16383) * 16384))
> bitAnd: 16r0FFFFFFF
>
>
> which will need some more bits.
>
>
> IMO it's not clear that SmallInteger>>identityHash should be
> implemented that way. Finding a permutation of the small
> integers that also behaves like a good quality hash function and
> evaluates quickly (in significantly less time and complexity
> than, say, Bob Jenkins' lookup3) would be a really interesting
> research project. I don't know if it's possible. If no such
> thing exists, then getting at least some rough idea of what's
> the minimum necessary complexity for such hash functions would
> be valuable.
>
> Looking at hashMultiply as a non-identity hash function, one
> would start having problems when significantly more than 2^28
> objects are stored in a single hashed collection. 2^28 objects
> with e.g. 12 bytes per header and a minimum of one instance
> variable (so the hash value isn't a instance-constant) stored in
> a hashed collection requires more than 4gb, so that is clearly a
> 64 bit image problem. In 64 bits, 2^28 objects with e.g. 16
> bytes per header and a minimum of one instance variable each is
> already 6gb, and a minimum of 8gb with the hashed collection itself.
>
> Because of these figures, I'd think improving the implementation
> of hashed collections takes priority over adding more
> non-identity hash function bits (as long as the existing hash
> values are of good quality).
>
> Did you look at the Hash Analysis Tool I wrote? It's in the
> Cincom public Store repository. It comes in two bundles: Hash
> Analysis Tool, and Hash Analysis Tool - Extensions. With
> everything loaded, the tool comes with 300+ hash functions and
> 100+ data sets out of the box. The code is MIT.
>
> I had a look at how it was done in VisualWorks;
>
>
> The implementation of hashMultiply, yes. Note however that
> SmallInteger>>hash is ^self.
>
> hashMultiply
> "Multiply the receiver by 16r0019660D mod 2^28
> without using large integer arithmetic for speed.
> The constant is a generator of the multiplicative
> subgroup of Z_2^30, see Knuth's TAOCP vol 2."
> <primitive: 1747>
> | low14Bits |
> low14Bits := self bitAnd: 16r3FFF.
> ^16384
> * (16r260D * (self bitShift: -14) + (16r0065 * low14Bits)
> bitAnd: 16r3FFF)
> + (16r260D * low14Bits) bitAnd: 16rFFFFFFF
>
> The hashing book version has:
>
> multiplication
> "Computes self times 1664525 mod 2^38 while avoiding
> overflow into a
> large integer by making the multiplication into two 14 bits
> chunks. Do
> not use any division or modulo operation."
> | lowBits highBits|
>
> lowBits := self bitAnd: 16r3FFF.
> highBits := self bitShift: -14.
> ^(lowBits * 16r260D)
> + (((lowBits * 16r0065) bitAnd: 16r3FFF) bitShift: 14)
> + (((highBits * 16r260D) bitAnd: 16r3FFF) bitShift: 14)
> bitAnd: 16rFFFFFFF
>
> So, 16384 * is the same as bitShift: 14 and it looks like
> done once,
> which may be better.
>
>
> It should be a primitive (or otherwise optimized somehow). At
> some point though that hash function was implemented for e.g.
> ByteArray in Squeak, I thought at that point the multiplication
> step was also implemented as a primitive?
>
> Also VW marks it as a primitive, which Pharo does not.
>
>
> In VW it is also a translated primitive, i.e. it's executed
> directly in the JIT without calling C.
>
> Keep in mind that the speed at which hash values are calculated
> is only part of the story. If the hash function quality is not
> great, or the hashed collection implementation is not efficient
> and induces collisions or other extra work, improving the
> efficiency of the hash functions won't do much. I think it's
> mentioned in the hash book (I'd have to check), but once I made
> a hash function 5x times slower to get better quality and the
> result was that a report that was taking 30 minutes took 90
> seconds instead (and hashing was gone from the profiler output).
>
> Would we gain
> some speed doing that? hashMultiply is used a lof for
> identity hashes.
>
> Bytecode has quite some work to do:
>
> 37 <70> self
> 38 <20> pushConstant: 16383
> 39 <BE> send: bitAnd:
> 40 <68> popIntoTemp: 0
> 41 <21> pushConstant: 9741
> 42 <10> pushTemp: 0
> 43 <B8> send: *
> 44 <21> pushConstant: 9741
> 45 <70> self
> 46 <22> pushConstant: -14
> 47 <BC> send: bitShift:
> 48 <B8> send: *
> 49 <23> pushConstant: 101
> 50 <10> pushTemp: 0
> 51 <B8> send: *
> 52 <B0> send: +
> 53 <20> pushConstant: 16383
> 54 <BE> send: bitAnd:
> 55 <24> pushConstant: 16384
> 56 <B8> send: *
> 57 <B0> send: +
> 58 <25> pushConstant: 268435455
> 59 <BE> send: bitAnd:
> 60 <7C> returnTop
>
>
> If this is a primitive instead, then you can also avoid the
> overflow into large integers and do the math with (basically)
>
> mov eax, smallInteger
> shr eax, numberOfTagBits
> mul eax, 1664525 "the multiplication that throws out the high bits"
> shl eax, 4 "throw out bits 29-32"
> shr eax, 4
> lea eax, [eax * 2^numberOfTagBits + smallIntegerTagBits]
>
> Please excuse trivial omissions in the above, it's written only
> for the sake of illustration (e.g. it looks like the 3 last
> instructions can be combined into two... lea followed by shr).
> Also, did you see the latency of integer multiplication
> instructions in modern x86 processors?...
>
> I ran some experiments timing things.
>
> It looks like that replacing 16384 * by bitShift:14 leads to
> a small
> gain, bitShift (primitive 17) being faster than * (primitive 9)
>
>
> Keep in mind those operations still have to check for overflow
> into large integers. In this case, large integers are not
> necessary.
>
> Andres.
>
> The bytecode is identical, except send: bitShift instead of
> send: *
>
> multiplication3
> | low |
>
> low := self bitAnd: 16383.
> ^(16r260D * low + ((16r260D * (self bitShift: -14) +
> (16r0065 * low)
> bitAnd: 16383) bitShift: 14))
> bitAnd: 16r0FFFFFFF
>
>
>
> [500000 timesRepeat: [ 15000 hashMultiply ]] timeToRun 12
> [500000 timesRepeat: [ 15000 multiplication ]] timeToRun 41
> (worse)
> [500000 timesRepeat: [ 15000 multiplication3 ]] timeToRun 10
> (better)
>
> It looks like correct for SmallInteger minVal to:
> SmallInteger maxVal
>
> Now, VW gives: [500000 timesRepeat: [ 15000 hashMultiply ]]
> timeToRun
> 1.149 milliseconds
>
> Definitely worth investigating the primitive thing, or some
> NB Asm as
> this is used about everywhere (Collections etc).
>
> Toughts?
>
> Phil
>
>
>
>
March 2, 2014
Re: [Pharo-dev] understanding postgresv2
by Yanni Chiu
Hmmm. Didn't remember that was there. IIRC, it was released under Squeak
Licence, via SqueakMap. Then when migrated to squeaksource.com, I
believe it was marked as MIT. I've lost track of where it's being
actively maintained, but please go ahead and remove or update the
copyright text in the comment to MIT.
--
Yanni
On 02/03/2014 11:36 AM, Tudor Girba wrote:
> Another thing I see in the comment of PGConnection is this:
> Copyright (c) 2001-2003 by Yanni Chiu. All Rights Reserved.
>
> Does anyone know the actual license?
>
> Doru
>
>
> On Sun, Mar 2, 2014 at 5:34 PM, Tudor Girba <tudor(a)tudorgirba.com
> <mailto:tudor@tudorgirba.com>> wrote:
>
> Hi,
>
> I am trying to understand how PostgresV2 is implemented because I
> would like to build some inspector support for it, and I encounter a
> couple of issues. In case anyone knows the answer, it would speed up
> my effort:
>
> - Why is result an instance variable in PGConnection? Making it a
> variable always returns the same object when executing a query and
> that is a bit of a pain.
>
> - Why does the PGResult have the possibility of holding multiple
> PGResultSets? When is it possible to have multiple at the same time?
> (when you execute a query, the result is being initialized)
>
> - When running something like
> connection execute: 'select * from ...'
> the PGResultSet already includes all rows of the query. Is it not
> possible to have a stream-like functionality in which the actual
> rows are retrieved only on demand? (a similar functionality exists
> in DBXTalk)
>
> - What is the difference between PGAsciiRow and PGDataRow?
>
> Doru
>
> --
> www.tudorgirba.com <http://www.tudorgirba.com>
>
> "Every thing has its own flow"
>
>
>
>
> --
> www.tudorgirba.com <http://www.tudorgirba.com>
>
> "Every thing has its own flow"
March 2, 2014
Re: [Pharo-dev] hashMultiply
by phil@highoctane.be
I have found the tools in the Cincom public repo.
Now, I have to see how this works in VW, I am not that proficient with it.
Phil
On Wed, Feb 26, 2014 at 8:55 AM, phil(a)highoctane.be <phil(a)highoctane.be>wrote:
> Andres,
>
> Thanks for the insights.
>
> hash quality is indeed an key factor. At least, my mind is somewhat
> grasping this hashing field a bit better.
>
> I'll have a look at the tools, I haven't used them yet.
>
> And a shot at the ASM version with NativeBoost in Pharo.
>
> For what occurs in modern CPUs, well, no. Surprising to see that a mul
> would be faster than a shr or shl. How comes?
> I used to be ok with these things when I was writing demoscene code a
> loong time ago but I'd need an extremely serious refresh.
>
> As a side note, there is a huge uptake in the BigData/mapreduce/hadoop
> environment where Smalltalk is sorely absent. Scala seems to fill the void
> on the JVM.
> There hashing is quite key, to remap all of the mapping phase results to
> the reduce nodes. I am surprised to see that Smalltalk vendors haven't
> jumped in that space.
>
> Phil
>
>
> On Wed, Feb 26, 2014 at 2:04 AM, Andres Valloud <
> avalloud(a)smalltalk.comcastbiz.net> wrote:
>
>> Hello...
>>
>> On 2/25/14 1:17 , phil(a)highoctane.be wrote:
>>
>>> I am currently reading through the Hashing in Smalltalk book
>>> (http://www.lulu.com/shop/andres-valloud/hashing-in-
>>> smalltalk-theory-and-practice/paperback/product-3788892.html)
>>> and, my head hurting notwithstanding, there are indeed a ton of gems in
>>> this system. As he mentions, doing the exercises brings a lot of extra
>>> :-)
>>>
>>
>> :) thank you.
>>
>> When going to 64-bit, and with the new ObjectMemory scheme, I guess a
>>> couple of identity hashing functions will come under scrutiny.
>>>
>>> e.g.
>>>
>>> SmallInteger>>hashMultiply
>>> | low |
>>>
>>> low := self bitAnd: 16383.
>>> ^(16r260D * low + ((16r260D * (self bitShift: -14) + (16r0065 * low)
>>> bitAnd: 16383) * 16384))
>>> bitAnd: 16r0FFFFFFF
>>>
>>>
>>> which will need some more bits.
>>>
>>
>> IMO it's not clear that SmallInteger>>identityHash should be implemented
>> that way. Finding a permutation of the small integers that also behaves
>> like a good quality hash function and evaluates quickly (in significantly
>> less time and complexity than, say, Bob Jenkins' lookup3) would be a really
>> interesting research project. I don't know if it's possible. If no such
>> thing exists, then getting at least some rough idea of what's the minimum
>> necessary complexity for such hash functions would be valuable.
>>
>> Looking at hashMultiply as a non-identity hash function, one would start
>> having problems when significantly more than 2^28 objects are stored in a
>> single hashed collection. 2^28 objects with e.g. 12 bytes per header and a
>> minimum of one instance variable (so the hash value isn't a
>> instance-constant) stored in a hashed collection requires more than 4gb, so
>> that is clearly a 64 bit image problem. In 64 bits, 2^28 objects with e.g.
>> 16 bytes per header and a minimum of one instance variable each is already
>> 6gb, and a minimum of 8gb with the hashed collection itself.
>>
>> Because of these figures, I'd think improving the implementation of
>> hashed collections takes priority over adding more non-identity hash
>> function bits (as long as the existing hash values are of good quality).
>>
>> Did you look at the Hash Analysis Tool I wrote? It's in the Cincom
>> public Store repository. It comes in two bundles: Hash Analysis Tool, and
>> Hash Analysis Tool - Extensions. With everything loaded, the tool comes
>> with 300+ hash functions and 100+ data sets out of the box. The code is
>> MIT.
>>
>> I had a look at how it was done in VisualWorks;
>>>
>>
>> The implementation of hashMultiply, yes. Note however that
>> SmallInteger>>hash is ^self.
>>
>> hashMultiply
>>> "Multiply the receiver by 16r0019660D mod 2^28
>>> without using large integer arithmetic for speed.
>>> The constant is a generator of the multiplicative
>>> subgroup of Z_2^30, see Knuth's TAOCP vol 2."
>>> <primitive: 1747>
>>> | low14Bits |
>>> low14Bits := self bitAnd: 16r3FFF.
>>> ^16384
>>> * (16r260D * (self bitShift: -14) + (16r0065 * low14Bits) bitAnd:
>>> 16r3FFF)
>>> + (16r260D * low14Bits) bitAnd: 16rFFFFFFF
>>>
>>> The hashing book version has:
>>>
>>> multiplication
>>> "Computes self times 1664525 mod 2^38 while avoiding overflow into a
>>> large integer by making the multiplication into two 14 bits chunks. Do
>>> not use any division or modulo operation."
>>> | lowBits highBits|
>>>
>>> lowBits := self bitAnd: 16r3FFF.
>>> highBits := self bitShift: -14.
>>> ^(lowBits * 16r260D)
>>> + (((lowBits * 16r0065) bitAnd: 16r3FFF) bitShift: 14)
>>> + (((highBits * 16r260D) bitAnd: 16r3FFF) bitShift: 14)
>>> bitAnd: 16rFFFFFFF
>>>
>>> So, 16384 * is the same as bitShift: 14 and it looks like done once,
>>> which may be better.
>>>
>>
>> It should be a primitive (or otherwise optimized somehow). At some point
>> though that hash function was implemented for e.g. ByteArray in Squeak, I
>> thought at that point the multiplication step was also implemented as a
>> primitive?
>>
>> Also VW marks it as a primitive, which Pharo does not.
>>>
>>
>> In VW it is also a translated primitive, i.e. it's executed directly in
>> the JIT without calling C.
>>
>> Keep in mind that the speed at which hash values are calculated is only
>> part of the story. If the hash function quality is not great, or the
>> hashed collection implementation is not efficient and induces collisions or
>> other extra work, improving the efficiency of the hash functions won't do
>> much. I think it's mentioned in the hash book (I'd have to check), but
>> once I made a hash function 5x times slower to get better quality and the
>> result was that a report that was taking 30 minutes took 90 seconds instead
>> (and hashing was gone from the profiler output).
>>
>> Would we gain
>>> some speed doing that? hashMultiply is used a lof for identity hashes.
>>>
>>> Bytecode has quite some work to do:
>>>
>>> 37 <70> self
>>> 38 <20> pushConstant: 16383
>>> 39 <BE> send: bitAnd:
>>> 40 <68> popIntoTemp: 0
>>> 41 <21> pushConstant: 9741
>>> 42 <10> pushTemp: 0
>>> 43 <B8> send: *
>>> 44 <21> pushConstant: 9741
>>> 45 <70> self
>>> 46 <22> pushConstant: -14
>>> 47 <BC> send: bitShift:
>>> 48 <B8> send: *
>>> 49 <23> pushConstant: 101
>>> 50 <10> pushTemp: 0
>>> 51 <B8> send: *
>>> 52 <B0> send: +
>>> 53 <20> pushConstant: 16383
>>> 54 <BE> send: bitAnd:
>>> 55 <24> pushConstant: 16384
>>> 56 <B8> send: *
>>> 57 <B0> send: +
>>> 58 <25> pushConstant: 268435455
>>> 59 <BE> send: bitAnd:
>>> 60 <7C> returnTop
>>>
>>
>> If this is a primitive instead, then you can also avoid the overflow into
>> large integers and do the math with (basically)
>>
>> mov eax, smallInteger
>> shr eax, numberOfTagBits
>> mul eax, 1664525 "the multiplication that throws out the high bits"
>> shl eax, 4 "throw out bits 29-32"
>> shr eax, 4
>> lea eax, [eax * 2^numberOfTagBits + smallIntegerTagBits]
>>
>> Please excuse trivial omissions in the above, it's written only for the
>> sake of illustration (e.g. it looks like the 3 last instructions can be
>> combined into two... lea followed by shr). Also, did you see the latency
>> of integer multiplication instructions in modern x86 processors?...
>>
>> I ran some experiments timing things.
>>>
>>> It looks like that replacing 16384 * by bitShift:14 leads to a small
>>> gain, bitShift (primitive 17) being faster than * (primitive 9)
>>>
>>
>> Keep in mind those operations still have to check for overflow into large
>> integers. In this case, large integers are not necessary.
>>
>> Andres.
>>
>> The bytecode is identical, except send: bitShift instead of send: *
>>>
>>> multiplication3
>>> | low |
>>>
>>> low := self bitAnd: 16383.
>>> ^(16r260D * low + ((16r260D * (self bitShift: -14) + (16r0065 * low)
>>> bitAnd: 16383) bitShift: 14))
>>> bitAnd: 16r0FFFFFFF
>>>
>>>
>>>
>>> [500000 timesRepeat: [ 15000 hashMultiply ]] timeToRun 12
>>> [500000 timesRepeat: [ 15000 multiplication ]] timeToRun 41 (worse)
>>> [500000 timesRepeat: [ 15000 multiplication3 ]] timeToRun 10 (better)
>>>
>>> It looks like correct for SmallInteger minVal to: SmallInteger maxVal
>>>
>>> Now, VW gives: [500000 timesRepeat: [ 15000 hashMultiply ]] timeToRun
>>> 1.149 milliseconds
>>>
>>> Definitely worth investigating the primitive thing, or some NB Asm as
>>> this is used about everywhere (Collections etc).
>>>
>>> Toughts?
>>>
>>> Phil
>>>
>>
>>
>
March 2, 2014
Re: [Pharo-dev] Versioneer docs?
by Christophe Demarey
Hi Phil,
Le 1 mars 2014 à 16:50, phil(a)highoctane.be a écrit :
> I am now porting my code to 3.0
>
> I am using Versioneer to look at my configuration as when I do load it in 3.0, it seems that there are some duplicate packages coming in my package-cache and I want to remove these dupes. (Mostly Seaside related).
>
> Versioneer is very helpful in representing the configuration.
>
> Now, is there any doc about its usage?
You can find some documentation here: http://chercheurs.lille.inria.fr/~demarey/Tech/Versionner
Don't hesitate to give feedback.
Versionner gets better with feedbacks I got latest weeks.
Maybe you already know but you can use the record directive with Metacello to see what will be loaded, i.e. resolved dependencies from the configuration: https://github.com/dalehenrich/metacello-work/blob/master/docs/MetacelloScr….
It may help to debug.
Best regards,
Christophe.
March 2, 2014
Re: [Pharo-dev] [Bug] IdentitySet>>size
by Max Leske
On 02.03.2014, at 22:55, Andres Valloud <avalloud(a)smalltalk.comcastbiz.net> wrote:
> So, just out of curiosity, how does the IdentitySet get "damagedâ?
Thatâs what Iâd like to know too :)
Nicolai posted an update to the issue which should shed some light on the problem. But at the moment I have no clue. Iâm hoping that somebody has worked with those tests and knows something.
>
> On 3/2/14 11:28 , Max Leske wrote:
>>
>> On 02.03.2014, at 20:12, Andres Valloud <avalloud(a)smalltalk.comcastbiz.net> wrote:
>>
>>> So it seems the problem is with Fuel rather than IdentitySet, no?
>>
>> No. Asking an IdentitySet for its size is not reliable. That has nothing to do with Fuel. #size is especially important in hashed collections where the size is stored in a variable and thus allows fast access the size without the need to visit every element in the collection for its computation.
>>
>>>
>>> On 3/2/14 10:47 , Max Leske wrote:
>>>> During serialization the IdentitySet size is stored and later its objects. During that step, #do:
>>>
>>
>>
>>
>
March 2, 2014