Pharo-users
By thread
pharo-users@lists.pharo.org
By month
Messages by month
- ----- 2026 -----
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
October 2019
- 74 participants
- 329 messages
Re: [Pharo-users] How to zip a WideString
by PBKResearch
Richard
I don't think so. The case being considered for my problem is the compression of a ByteArray produced by applying #utf8Encoded to a WideString, but it extends to any other form of ByteArray. If you substitute ByteArray for SomeClass in your examples, I think you will see why the chosen interface was used.
Peter Kenny
-----Original Message-----
From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of Richard O'Keefe
Sent: 03 October 2019 23:08
To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
Subject: Re: [Pharo-users] How to zip a WideString
The interface should surely be
SomeClass
methods for: 'compression'
zipped "return a byte array"
class methods for: 'decompression'
unzip: aByteArray "return an instance of SomeClass"
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by Richard O'Keefe
The interface should surely be
SomeClass
methods for: 'compression'
zipped "return a byte array"
class methods for: 'decompression'
unzip: aByteArray "return an instance of SomeClass"
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by Sven Van Caekenberghe
> On 3 Oct 2019, at 18:25, Sean P. DeNigris <sean(a)clipperadams.com> wrote:
>
> Peter Kenny wrote
>> Just 5 hours from when I raised the question, there is a solution in place
>> for everyone. This group is amazing!
>
> Indeed. Bravo, Sven and all of our other contributors.
Thanks, but I must say that it was Tomo's snippet that showed me how to use the gzip streams binary to binary as I thought that was not possible.
And we need user's questions to start acting, so thank you Peter as well.
> I can't resist mentioning that the fix would almost certainly have taken
> significantly longer a few years ago due to the awkward issue/contribution
> infrastructure. Esteban and others' tireless work to get git/GH support up
> and running are at the heart of all these quick turnaround stories. Keep
> going!!
>
> p.s. it takes serious courage to port the issue tracker of a project this
> size and reach twice in a few years
Indeed, it had been a while since I did a PR and I must say it went super smooth.
Also, Pharo 8 feels quite snappy.
> -----
> Cheers,
> Sean
> --
> Sent from: http://forum.world.st/Pharo-Smalltalk-Users-f1310670.html
>
Oct. 3, 2019
Re: [Pharo-users] [ANN] HeySql - a mini db-orm for Postgres
by Sean P. DeNigris
Petter Egesund wrote
> Hi, nice to meet you all :)
Code submission is a wonderful way to meet ha ha. Thanks!
-----
Cheers,
Sean
--
Sent from: http://forum.world.st/Pharo-Smalltalk-Users-f1310670.html
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by Sean P. DeNigris
Peter Kenny wrote
> Just 5 hours from when I raised the question, there is a solution in place
> for everyone. This group is amazing!
Indeed. Bravo, Sven and all of our other contributors.
I can't resist mentioning that the fix would almost certainly have taken
significantly longer a few years ago due to the awkward issue/contribution
infrastructure. Esteban and others' tireless work to get git/GH support up
and running are at the heart of all these quick turnaround stories. Keep
going!!
p.s. it takes serious courage to port the issue tracker of a project this
size and reach twice in a few years
-----
Cheers,
Sean
--
Sent from: http://forum.world.st/Pharo-Smalltalk-Users-f1310670.html
Oct. 3, 2019
[ANN] HeySql - a mini db-orm for Postgres
by Petter Egesund
Hi, nice to meet you all :)
As my first Smalltalk attempt I have written a mini orm for Postgres, based
on the P3-library.
Some features:
- Closely connected to sql
- Easy to use
- Runtime compiled methods based on server side statements
- Generated methods for insert, update on method objects
- Generate tables
It can be found here:
https://github.com/pegesund/heysql
I did also write something about the code, Pharo and what I think on my
kind-of-blog, if anyone is interested:
https://ramblings.work/posts/2019-02-10-heysql.html
Best regards,
Petter Egesund
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by PBKResearch
Thanks Sven. Just 5 hours from when I raised the question, there is a solution in place for everyone. This group is amazing!
-----Original Message-----
From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of Sven Van Caekenberghe
Sent: 03 October 2019 15:28
To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
Subject: Re: [Pharo-users] How to zip a WideString
https://github.com/pharo-project/pharo/pull/4812
> On 3 Oct 2019, at 14:05, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
>
> https://github.com/pharo-project/pharo/issues/4806
>
> PR will follow
>
>> On 3 Oct 2019, at 13:49, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>>
>> Sven, Tomo
>>
>> Thanks for this discussion. I shall bear in mind Sven's proposed extension to ByteArray - this is exactly the sort of neater solution I was hoping for. Any chance this might make it into standard Pharo (perhaps inP8)?
>>
>> Peter Kenny
>>
>> -----Original Message-----
>> From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of
>> Tomohiro Oda
>> Sent: 03 October 2019 12:22
>> To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
>> Subject: Re: [Pharo-users] How to zip a WideString
>>
>> Sven,
>>
>> Yes, ByteArray>>zipped/unzipped are simple, neat and intuitive way of zipping/unzipping binary data.
>> I also love the new idioms. They look clean and concise.
>>
>> Best Regards,
>> ---
>> tomo
>>
>> 2019å¹´10æ3æ¥(æ¨) 20:14 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>>
>>> Actually, thinking about this a bit more, why not add #zipped #unzipped to ByteArray ?
>>>
>>>
>>> ByteArray>>#zipped
>>> "Return a GZIP compressed version of the receiver as a ByteArray"
>>>
>>> ^ ByteArray streamContents: [ :out |
>>> (GZipWriteStream on: out) nextPutAll: self; close ]
>>>
>>> ByteArray>>#unzipped
>>> "Assuming the receiver contains GZIP encoded data, return the
>>> decompressed data as a ByteArray"
>>>
>>> ^ (GZipReadStream on: self) upToEnd
>>>
>>>
>>> The original oneliner then becomes
>>>
>>> 'string' utf8Encoded zipped.
>>>
>>> and
>>>
>>> data unzipped utf8Decoded
>>>
>>> which is pretty clear, simple and intention-revealing, IMHO.
>>>
>>>> On 3 Oct 2019, at 13:04, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
>>>>
>>>> Hi Tomo,
>>>>
>>>> Indeed, I stand corrected, it does indeed seem possible to use the existing gzip classes to work from bytes to bytes, this works fine:
>>>>
>>>> data := ByteArray streamContents: [ :out | (GZipWriteStream on: out) nextPutAll: 'foo 10 â¬' utf8Encoded; close ].
>>>>
>>>> (GZipReadStream on: data) upToEnd utf8Decoded.
>>>>
>>>> Now regarding the encoding option, I am not so sure that is really necessary (though nice to have). Why would anyone use anything except UTF8 (today).
>>>>
>>>> Thanks again for the correction !
>>>>
>>>> Sven
>>>>
>>>>> On 3 Oct 2019, at 12:41, Tomohiro Oda <tomohiro.tomo.oda(a)gmail.com> wrote:
>>>>>
>>>>> Peter and Sven,
>>>>>
>>>>> zip API from string to string works fine except that aWideString
>>>>> zipped generates malformed zip string.
>>>>> I think it might be a good guidance to define
>>>>> String>>zippedWithEncoding: and ByteArray>>unzippedWithEncoding: .
>>>>> Such as
>>>>> String>>zippedWithEncoding: encoder
>>>>> zippedWithEncoding: encoder
>>>>> ^ ByteArray
>>>>> streamContents: [ :stream |
>>>>> | gzstream |
>>>>> gzstream := GZipWriteStream on: stream.
>>>>> encoder
>>>>> next: self size
>>>>> putAll: self
>>>>> startingAt: 1
>>>>> toStream: gzstream.
>>>>> gzstream close ]
>>>>>
>>>>> and ByteArray>>unzippedWithEncoding: encoder
>>>>> unzippedWithEncoding: encoder
>>>>> | byteStream |
>>>>> byteStream := GZipReadStream on: self.
>>>>> ^ String
>>>>> streamContents: [ :stream |
>>>>> [ byteStream atEnd ]
>>>>> whileFalse: [ stream nextPut: (encoder nextFromStream:
>>>>> byteStream) ] ]
>>>>>
>>>>> Then, you can write something like zipped := yourLongWideString
>>>>> zippedWithEncoding: ZnCharacterEncoder utf8.
>>>>> unzipped := zipped unzippedWithEncoding: ZnCharacterEncoder utf8.
>>>>>
>>>>> This will not affect the existing zipped/unzipped API and you can
>>>>> handle other encodings.
>>>>> This zippedWithEncoding: generates a ByteArray, which is kind of
>>>>> conformant to the encoding API.
>>>>> And you don't have to create many intermediate byte arrays and byte strings.
>>>>>
>>>>> I hope this helps.
>>>>> ---
>>>>> tomo
>>>>>
>>>>> 2019/10/3(Thu) 18:56 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>>>>>
>>>>>> Hi Peter,
>>>>>>
>>>>>> About #zipped / #unzipped and the inflate / deflate classes: your observation is correct, these work from string to string, while clearly the compressed representation should be binary.
>>>>>>
>>>>>> The contents (input, what is inside the compressed data) can be anything, it is not necessarily a string (it could be an image, so also something binary). Only the creator of the compressed data knows, you cannot assume to know in general.
>>>>>>
>>>>>> It would be possible (and it would be very nice) to change this, however that will have serious impact on users (as the contract changes).
>>>>>>
>>>>>> About your use case: why would your DB not be capable of storing large strings ? A good DB should be capable of storing any kind of string (full unicode) efficiently.
>>>>>>
>>>>>> What DB and what sizes are we talking about ?
>>>>>>
>>>>>> Sven
>>>>>>
>>>>>>> On 3 Oct 2019, at 11:06, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>>>>>>>
>>>>>>> Hello
>>>>>>>
>>>>>>> I have a problem with text storage, to which I seem to have found a solution, but itâs a bit clumsy-looking. I would be grateful for confirmation that (a) there is no neater solution, (b) I can rely on this to work â I only know that it works in a few test cases.
>>>>>>>
>>>>>>> I need to store a large number of text strings in a database. To avoid the database files becoming too large, I am thinking of zipping the strings, or at least the less frequently accessed ones. Depending on the source, some of the strings will be instances of ByteString, some of WideString (because they contain characters not representable in one byte). Storing a WideString uncompressed seems to occupy 4 bytes per character, so I decided, before thinking of compression, to store the strings utf8Encoded, which yields a ByteArray. But zipped can only be applied to a String, not a ByteArray.
>>>>>>>
>>>>>>> So my proposed solution is:
>>>>>>>
>>>>>>> For compression: myZipString := myWideString utf8Encoded asString zipped.
>>>>>>> For decompression: myOutputString := myZipString unzipped asByteArray utf8Decoded.
>>>>>>>
>>>>>>> As I said, it works in all the cases I tried, whether WideString or not, but the chains of transformations look clunky somehow. Can anyone see a neater way of doing it? And can I rely on it working, especially when I am handling foreign texts with many multi-byte characters?
>>>>>>>
>>>>>>> Thanks in advance for any help.
>>>>>>>
>>>>>>> Peter Kenny
>>>>>>
>>>>>>
>>>>>
>>>>
>>>
>>>
>>
>>
>
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by Sven Van Caekenberghe
https://github.com/pharo-project/pharo/pull/4812
> On 3 Oct 2019, at 14:05, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
>
> https://github.com/pharo-project/pharo/issues/4806
>
> PR will follow
>
>> On 3 Oct 2019, at 13:49, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>>
>> Sven, Tomo
>>
>> Thanks for this discussion. I shall bear in mind Sven's proposed extension to ByteArray - this is exactly the sort of neater solution I was hoping for. Any chance this might make it into standard Pharo (perhaps inP8)?
>>
>> Peter Kenny
>>
>> -----Original Message-----
>> From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of Tomohiro Oda
>> Sent: 03 October 2019 12:22
>> To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
>> Subject: Re: [Pharo-users] How to zip a WideString
>>
>> Sven,
>>
>> Yes, ByteArray>>zipped/unzipped are simple, neat and intuitive way of zipping/unzipping binary data.
>> I also love the new idioms. They look clean and concise.
>>
>> Best Regards,
>> ---
>> tomo
>>
>> 2019å¹´10æ3æ¥(æ¨) 20:14 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>>
>>> Actually, thinking about this a bit more, why not add #zipped #unzipped to ByteArray ?
>>>
>>>
>>> ByteArray>>#zipped
>>> "Return a GZIP compressed version of the receiver as a ByteArray"
>>>
>>> ^ ByteArray streamContents: [ :out |
>>> (GZipWriteStream on: out) nextPutAll: self; close ]
>>>
>>> ByteArray>>#unzipped
>>> "Assuming the receiver contains GZIP encoded data,
>>> return the decompressed data as a ByteArray"
>>>
>>> ^ (GZipReadStream on: self) upToEnd
>>>
>>>
>>> The original oneliner then becomes
>>>
>>> 'string' utf8Encoded zipped.
>>>
>>> and
>>>
>>> data unzipped utf8Decoded
>>>
>>> which is pretty clear, simple and intention-revealing, IMHO.
>>>
>>>> On 3 Oct 2019, at 13:04, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
>>>>
>>>> Hi Tomo,
>>>>
>>>> Indeed, I stand corrected, it does indeed seem possible to use the existing gzip classes to work from bytes to bytes, this works fine:
>>>>
>>>> data := ByteArray streamContents: [ :out | (GZipWriteStream on: out) nextPutAll: 'foo 10 â¬' utf8Encoded; close ].
>>>>
>>>> (GZipReadStream on: data) upToEnd utf8Decoded.
>>>>
>>>> Now regarding the encoding option, I am not so sure that is really necessary (though nice to have). Why would anyone use anything except UTF8 (today).
>>>>
>>>> Thanks again for the correction !
>>>>
>>>> Sven
>>>>
>>>>> On 3 Oct 2019, at 12:41, Tomohiro Oda <tomohiro.tomo.oda(a)gmail.com> wrote:
>>>>>
>>>>> Peter and Sven,
>>>>>
>>>>> zip API from string to string works fine except that aWideString
>>>>> zipped generates malformed zip string.
>>>>> I think it might be a good guidance to define
>>>>> String>>zippedWithEncoding: and ByteArray>>unzippedWithEncoding: .
>>>>> Such as
>>>>> String>>zippedWithEncoding: encoder
>>>>> zippedWithEncoding: encoder
>>>>> ^ ByteArray
>>>>> streamContents: [ :stream |
>>>>> | gzstream |
>>>>> gzstream := GZipWriteStream on: stream.
>>>>> encoder
>>>>> next: self size
>>>>> putAll: self
>>>>> startingAt: 1
>>>>> toStream: gzstream.
>>>>> gzstream close ]
>>>>>
>>>>> and ByteArray>>unzippedWithEncoding: encoder
>>>>> unzippedWithEncoding: encoder
>>>>> | byteStream |
>>>>> byteStream := GZipReadStream on: self.
>>>>> ^ String
>>>>> streamContents: [ :stream |
>>>>> [ byteStream atEnd ]
>>>>> whileFalse: [ stream nextPut: (encoder nextFromStream:
>>>>> byteStream) ] ]
>>>>>
>>>>> Then, you can write something like
>>>>> zipped := yourLongWideString zippedWithEncoding: ZnCharacterEncoder utf8.
>>>>> unzipped := zipped unzippedWithEncoding: ZnCharacterEncoder utf8.
>>>>>
>>>>> This will not affect the existing zipped/unzipped API and you can
>>>>> handle other encodings.
>>>>> This zippedWithEncoding: generates a ByteArray, which is kind of
>>>>> conformant to the encoding API.
>>>>> And you don't have to create many intermediate byte arrays and byte strings.
>>>>>
>>>>> I hope this helps.
>>>>> ---
>>>>> tomo
>>>>>
>>>>> 2019/10/3(Thu) 18:56 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>>>>>
>>>>>> Hi Peter,
>>>>>>
>>>>>> About #zipped / #unzipped and the inflate / deflate classes: your observation is correct, these work from string to string, while clearly the compressed representation should be binary.
>>>>>>
>>>>>> The contents (input, what is inside the compressed data) can be anything, it is not necessarily a string (it could be an image, so also something binary). Only the creator of the compressed data knows, you cannot assume to know in general.
>>>>>>
>>>>>> It would be possible (and it would be very nice) to change this, however that will have serious impact on users (as the contract changes).
>>>>>>
>>>>>> About your use case: why would your DB not be capable of storing large strings ? A good DB should be capable of storing any kind of string (full unicode) efficiently.
>>>>>>
>>>>>> What DB and what sizes are we talking about ?
>>>>>>
>>>>>> Sven
>>>>>>
>>>>>>> On 3 Oct 2019, at 11:06, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>>>>>>>
>>>>>>> Hello
>>>>>>>
>>>>>>> I have a problem with text storage, to which I seem to have found a solution, but itâs a bit clumsy-looking. I would be grateful for confirmation that (a) there is no neater solution, (b) I can rely on this to work â I only know that it works in a few test cases.
>>>>>>>
>>>>>>> I need to store a large number of text strings in a database. To avoid the database files becoming too large, I am thinking of zipping the strings, or at least the less frequently accessed ones. Depending on the source, some of the strings will be instances of ByteString, some of WideString (because they contain characters not representable in one byte). Storing a WideString uncompressed seems to occupy 4 bytes per character, so I decided, before thinking of compression, to store the strings utf8Encoded, which yields a ByteArray. But zipped can only be applied to a String, not a ByteArray.
>>>>>>>
>>>>>>> So my proposed solution is:
>>>>>>>
>>>>>>> For compression: myZipString := myWideString utf8Encoded asString zipped.
>>>>>>> For decompression: myOutputString := myZipString unzipped asByteArray utf8Decoded.
>>>>>>>
>>>>>>> As I said, it works in all the cases I tried, whether WideString or not, but the chains of transformations look clunky somehow. Can anyone see a neater way of doing it? And can I rely on it working, especially when I am handling foreign texts with many multi-byte characters?
>>>>>>>
>>>>>>> Thanks in advance for any help.
>>>>>>>
>>>>>>> Peter Kenny
>>>>>>
>>>>>>
>>>>>
>>>>
>>>
>>>
>>
>>
>
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by Sven Van Caekenberghe
https://github.com/pharo-project/pharo/issues/4806
PR will follow
> On 3 Oct 2019, at 13:49, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>
> Sven, Tomo
>
> Thanks for this discussion. I shall bear in mind Sven's proposed extension to ByteArray - this is exactly the sort of neater solution I was hoping for. Any chance this might make it into standard Pharo (perhaps inP8)?
>
> Peter Kenny
>
> -----Original Message-----
> From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of Tomohiro Oda
> Sent: 03 October 2019 12:22
> To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
> Subject: Re: [Pharo-users] How to zip a WideString
>
> Sven,
>
> Yes, ByteArray>>zipped/unzipped are simple, neat and intuitive way of zipping/unzipping binary data.
> I also love the new idioms. They look clean and concise.
>
> Best Regards,
> ---
> tomo
>
> 2019å¹´10æ3æ¥(æ¨) 20:14 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>
>> Actually, thinking about this a bit more, why not add #zipped #unzipped to ByteArray ?
>>
>>
>> ByteArray>>#zipped
>> "Return a GZIP compressed version of the receiver as a ByteArray"
>>
>> ^ ByteArray streamContents: [ :out |
>> (GZipWriteStream on: out) nextPutAll: self; close ]
>>
>> ByteArray>>#unzipped
>> "Assuming the receiver contains GZIP encoded data,
>> return the decompressed data as a ByteArray"
>>
>> ^ (GZipReadStream on: self) upToEnd
>>
>>
>> The original oneliner then becomes
>>
>> 'string' utf8Encoded zipped.
>>
>> and
>>
>> data unzipped utf8Decoded
>>
>> which is pretty clear, simple and intention-revealing, IMHO.
>>
>>> On 3 Oct 2019, at 13:04, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
>>>
>>> Hi Tomo,
>>>
>>> Indeed, I stand corrected, it does indeed seem possible to use the existing gzip classes to work from bytes to bytes, this works fine:
>>>
>>> data := ByteArray streamContents: [ :out | (GZipWriteStream on: out) nextPutAll: 'foo 10 â¬' utf8Encoded; close ].
>>>
>>> (GZipReadStream on: data) upToEnd utf8Decoded.
>>>
>>> Now regarding the encoding option, I am not so sure that is really necessary (though nice to have). Why would anyone use anything except UTF8 (today).
>>>
>>> Thanks again for the correction !
>>>
>>> Sven
>>>
>>>> On 3 Oct 2019, at 12:41, Tomohiro Oda <tomohiro.tomo.oda(a)gmail.com> wrote:
>>>>
>>>> Peter and Sven,
>>>>
>>>> zip API from string to string works fine except that aWideString
>>>> zipped generates malformed zip string.
>>>> I think it might be a good guidance to define
>>>> String>>zippedWithEncoding: and ByteArray>>unzippedWithEncoding: .
>>>> Such as
>>>> String>>zippedWithEncoding: encoder
>>>> zippedWithEncoding: encoder
>>>> ^ ByteArray
>>>> streamContents: [ :stream |
>>>> | gzstream |
>>>> gzstream := GZipWriteStream on: stream.
>>>> encoder
>>>> next: self size
>>>> putAll: self
>>>> startingAt: 1
>>>> toStream: gzstream.
>>>> gzstream close ]
>>>>
>>>> and ByteArray>>unzippedWithEncoding: encoder
>>>> unzippedWithEncoding: encoder
>>>> | byteStream |
>>>> byteStream := GZipReadStream on: self.
>>>> ^ String
>>>> streamContents: [ :stream |
>>>> [ byteStream atEnd ]
>>>> whileFalse: [ stream nextPut: (encoder nextFromStream:
>>>> byteStream) ] ]
>>>>
>>>> Then, you can write something like
>>>> zipped := yourLongWideString zippedWithEncoding: ZnCharacterEncoder utf8.
>>>> unzipped := zipped unzippedWithEncoding: ZnCharacterEncoder utf8.
>>>>
>>>> This will not affect the existing zipped/unzipped API and you can
>>>> handle other encodings.
>>>> This zippedWithEncoding: generates a ByteArray, which is kind of
>>>> conformant to the encoding API.
>>>> And you don't have to create many intermediate byte arrays and byte strings.
>>>>
>>>> I hope this helps.
>>>> ---
>>>> tomo
>>>>
>>>> 2019/10/3(Thu) 18:56 Sven Van Caekenberghe <sven(a)stfx.eu>:
>>>>>
>>>>> Hi Peter,
>>>>>
>>>>> About #zipped / #unzipped and the inflate / deflate classes: your observation is correct, these work from string to string, while clearly the compressed representation should be binary.
>>>>>
>>>>> The contents (input, what is inside the compressed data) can be anything, it is not necessarily a string (it could be an image, so also something binary). Only the creator of the compressed data knows, you cannot assume to know in general.
>>>>>
>>>>> It would be possible (and it would be very nice) to change this, however that will have serious impact on users (as the contract changes).
>>>>>
>>>>> About your use case: why would your DB not be capable of storing large strings ? A good DB should be capable of storing any kind of string (full unicode) efficiently.
>>>>>
>>>>> What DB and what sizes are we talking about ?
>>>>>
>>>>> Sven
>>>>>
>>>>>> On 3 Oct 2019, at 11:06, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
>>>>>>
>>>>>> Hello
>>>>>>
>>>>>> I have a problem with text storage, to which I seem to have found a solution, but itâs a bit clumsy-looking. I would be grateful for confirmation that (a) there is no neater solution, (b) I can rely on this to work â I only know that it works in a few test cases.
>>>>>>
>>>>>> I need to store a large number of text strings in a database. To avoid the database files becoming too large, I am thinking of zipping the strings, or at least the less frequently accessed ones. Depending on the source, some of the strings will be instances of ByteString, some of WideString (because they contain characters not representable in one byte). Storing a WideString uncompressed seems to occupy 4 bytes per character, so I decided, before thinking of compression, to store the strings utf8Encoded, which yields a ByteArray. But zipped can only be applied to a String, not a ByteArray.
>>>>>>
>>>>>> So my proposed solution is:
>>>>>>
>>>>>> For compression: myZipString := myWideString utf8Encoded asString zipped.
>>>>>> For decompression: myOutputString := myZipString unzipped asByteArray utf8Decoded.
>>>>>>
>>>>>> As I said, it works in all the cases I tried, whether WideString or not, but the chains of transformations look clunky somehow. Can anyone see a neater way of doing it? And can I rely on it working, especially when I am handling foreign texts with many multi-byte characters?
>>>>>>
>>>>>> Thanks in advance for any help.
>>>>>>
>>>>>> Peter Kenny
>>>>>
>>>>>
>>>>
>>>
>>
>>
>
>
Oct. 3, 2019
Re: [Pharo-users] How to zip a WideString
by PBKResearch
Sven, Tomo
Thanks for this discussion. I shall bear in mind Sven's proposed extension to ByteArray - this is exactly the sort of neater solution I was hoping for. Any chance this might make it into standard Pharo (perhaps inP8)?
Peter Kenny
-----Original Message-----
From: Pharo-users <pharo-users-bounces(a)lists.pharo.org> On Behalf Of Tomohiro Oda
Sent: 03 October 2019 12:22
To: Any question about pharo is welcome <pharo-users(a)lists.pharo.org>
Subject: Re: [Pharo-users] How to zip a WideString
Sven,
Yes, ByteArray>>zipped/unzipped are simple, neat and intuitive way of zipping/unzipping binary data.
I also love the new idioms. They look clean and concise.
Best Regards,
---
tomo
2019å¹´10æ3æ¥(æ¨) 20:14 Sven Van Caekenberghe <sven(a)stfx.eu>:
>
> Actually, thinking about this a bit more, why not add #zipped #unzipped to ByteArray ?
>
>
> ByteArray>>#zipped
> "Return a GZIP compressed version of the receiver as a ByteArray"
>
> ^ ByteArray streamContents: [ :out |
> (GZipWriteStream on: out) nextPutAll: self; close ]
>
> ByteArray>>#unzipped
> "Assuming the receiver contains GZIP encoded data,
> return the decompressed data as a ByteArray"
>
> ^ (GZipReadStream on: self) upToEnd
>
>
> The original oneliner then becomes
>
> 'string' utf8Encoded zipped.
>
> and
>
> data unzipped utf8Decoded
>
> which is pretty clear, simple and intention-revealing, IMHO.
>
> > On 3 Oct 2019, at 13:04, Sven Van Caekenberghe <sven(a)stfx.eu> wrote:
> >
> > Hi Tomo,
> >
> > Indeed, I stand corrected, it does indeed seem possible to use the existing gzip classes to work from bytes to bytes, this works fine:
> >
> > data := ByteArray streamContents: [ :out | (GZipWriteStream on: out) nextPutAll: 'foo 10 â¬' utf8Encoded; close ].
> >
> > (GZipReadStream on: data) upToEnd utf8Decoded.
> >
> > Now regarding the encoding option, I am not so sure that is really necessary (though nice to have). Why would anyone use anything except UTF8 (today).
> >
> > Thanks again for the correction !
> >
> > Sven
> >
> >> On 3 Oct 2019, at 12:41, Tomohiro Oda <tomohiro.tomo.oda(a)gmail.com> wrote:
> >>
> >> Peter and Sven,
> >>
> >> zip API from string to string works fine except that aWideString
> >> zipped generates malformed zip string.
> >> I think it might be a good guidance to define
> >> String>>zippedWithEncoding: and ByteArray>>unzippedWithEncoding: .
> >> Such as
> >> String>>zippedWithEncoding: encoder
> >> zippedWithEncoding: encoder
> >> ^ ByteArray
> >> streamContents: [ :stream |
> >> | gzstream |
> >> gzstream := GZipWriteStream on: stream.
> >> encoder
> >> next: self size
> >> putAll: self
> >> startingAt: 1
> >> toStream: gzstream.
> >> gzstream close ]
> >>
> >> and ByteArray>>unzippedWithEncoding: encoder
> >> unzippedWithEncoding: encoder
> >> | byteStream |
> >> byteStream := GZipReadStream on: self.
> >> ^ String
> >> streamContents: [ :stream |
> >> [ byteStream atEnd ]
> >> whileFalse: [ stream nextPut: (encoder nextFromStream:
> >> byteStream) ] ]
> >>
> >> Then, you can write something like
> >> zipped := yourLongWideString zippedWithEncoding: ZnCharacterEncoder utf8.
> >> unzipped := zipped unzippedWithEncoding: ZnCharacterEncoder utf8.
> >>
> >> This will not affect the existing zipped/unzipped API and you can
> >> handle other encodings.
> >> This zippedWithEncoding: generates a ByteArray, which is kind of
> >> conformant to the encoding API.
> >> And you don't have to create many intermediate byte arrays and byte strings.
> >>
> >> I hope this helps.
> >> ---
> >> tomo
> >>
> >> 2019/10/3(Thu) 18:56 Sven Van Caekenberghe <sven(a)stfx.eu>:
> >>>
> >>> Hi Peter,
> >>>
> >>> About #zipped / #unzipped and the inflate / deflate classes: your observation is correct, these work from string to string, while clearly the compressed representation should be binary.
> >>>
> >>> The contents (input, what is inside the compressed data) can be anything, it is not necessarily a string (it could be an image, so also something binary). Only the creator of the compressed data knows, you cannot assume to know in general.
> >>>
> >>> It would be possible (and it would be very nice) to change this, however that will have serious impact on users (as the contract changes).
> >>>
> >>> About your use case: why would your DB not be capable of storing large strings ? A good DB should be capable of storing any kind of string (full unicode) efficiently.
> >>>
> >>> What DB and what sizes are we talking about ?
> >>>
> >>> Sven
> >>>
> >>>> On 3 Oct 2019, at 11:06, PBKResearch <peter(a)pbkresearch.co.uk> wrote:
> >>>>
> >>>> Hello
> >>>>
> >>>> I have a problem with text storage, to which I seem to have found a solution, but itâs a bit clumsy-looking. I would be grateful for confirmation that (a) there is no neater solution, (b) I can rely on this to work â I only know that it works in a few test cases.
> >>>>
> >>>> I need to store a large number of text strings in a database. To avoid the database files becoming too large, I am thinking of zipping the strings, or at least the less frequently accessed ones. Depending on the source, some of the strings will be instances of ByteString, some of WideString (because they contain characters not representable in one byte). Storing a WideString uncompressed seems to occupy 4 bytes per character, so I decided, before thinking of compression, to store the strings utf8Encoded, which yields a ByteArray. But zipped can only be applied to a String, not a ByteArray.
> >>>>
> >>>> So my proposed solution is:
> >>>>
> >>>> For compression: myZipString := myWideString utf8Encoded asString zipped.
> >>>> For decompression: myOutputString := myZipString unzipped asByteArray utf8Decoded.
> >>>>
> >>>> As I said, it works in all the cases I tried, whether WideString or not, but the chains of transformations look clunky somehow. Can anyone see a neater way of doing it? And can I rely on it working, especially when I am handling foreign texts with many multi-byte characters?
> >>>>
> >>>> Thanks in advance for any help.
> >>>>
> >>>> Peter Kenny
> >>>
> >>>
> >>
> >
>
>
Oct. 3, 2019