11:42
<annevk>
so only Mozilla has implemented ele.spellcheck and they do it as a boolean rather than enumerable attribute as the spec requires?
11:42
<jgraham>
annevk: Yes
11:43
<annevk>
I guess the spec wanted it to be in sync with .contentEditable
11:43
<jgraham>
hybi list fail - Thomas was blue not green. Henry was green. (and James was red)
11:43
<annevk>
maybe even on my request
11:44
<annevk>
jgraham, euh?!
11:44
<jgraham>
Yeah, I think the spec isn't going to happen
11:45
<jgraham>
annevk: Greg's extension of Hixie's transport analogy had a vehicle that is green, runs on rails, and answers to the name Thomas
11:45
<jgraham>
But Thomas was blue
11:46
<jgraham>
(possibly Thomas the Tank Engine is an Anglo-American curio)
11:47
<annevk>
Filed a bug on spellcheck.
11:47
<jgraham>
annevk: I already filed a bug indicating that throwing SYNTAX_ERR wasn't going to work
12:25
<MikeSmith>
annevk: I added an "HTML elements organized by function" section to the HtmlR doc -
12:25
<MikeSmith>
http://dev.w3.org/html5/markup/elements-by-function.html
12:25
<MikeSmith>
(I think you had suggested it should have one)
13:04
<annevk>
ah yeah
13:04
<erlehmann_>
annevk, I want to do some very light DOM manipulation in PHP and intend to use html5lib to get the DOM. Two question: First, is that a good choice? Second, what solution would you recommend to manipulate said DOM in PHP?
13:04
<annevk>
well, I suggested the main draft would be done in that way, but I suppose this works too
13:04
<annevk>
erlehmann_, I don't have experience with the PHP html5lib unfortunately
13:05
<annevk>
erlehmann, nor with PHP DOM manipulation :(
13:05
<annevk>
erlehmann, though overall that sounds like the best way if you plan on using PHP
13:05
<erlehmann>
just saw you as project owner on google code. that's the python part then, right?
13:06
<jgraham>
erlehmann: Well I think PHP isn't a good choice :)
13:06
<jgraham>
But if hat is a constraint then html5lib is a good choice for parsing the HTML
13:06
<jgraham>
as long as speed is not your main concern
13:07
<jgraham>
gsnedders and ezyang were mainly responsible for the PHP verson
13:07
<erlehmann>
jgraham, i agree wholeheartedly. when i applied for an internship in early 2008 and they asked me if i knew PHP, i told them why i hate it.
13:07
<annevk>
erlehmann, oh, yeah, I did the original Python version in part way back though jgraham knows and did more :)
13:07
<erlehmann>
but right now i have a gsoc project to finish.
13:08
<erlehmann>
and since i am writing a wordpress plugin … well ;)
13:08
<jgraham>
Ah
13:08
<jgraham>
You chose the wrong problem
13:08
<jgraham>
:)
13:09
<jgraham>
It is at least worth trying using PHP html5lib
13:09
<erlehmann>
probably. but all my friends are using wordpress.
13:09
<erlehmann>
and me too. though i will look into habari Really Soon Now [TM]
13:10
<jgraham>
That's PHP too, right?
13:10
<erlehmann>
yeah. i should probably bully my hoster into getting me some WSGI goodness so i can install a python-based imageboard instead of my boring old blog.
13:13
<erlehmann>
i'll look if PHP Simple HTML DOM Parser does it for me. i do not have that many edge cases and it looks nice and usable.
14:17
<gsnedders>
The PHP html5lib is really quite out of date
14:17
<gsnedders>
There's access to the libxml HTML parser from the DOM extension
19:11
<Hixie>
jgraham: yeah i considered saying gordon was green and thomas wasn't an eletric locomotive, but i figured that was maybe being pedantic about hte wrong thing :-)
19:12
<Hixie>
wait, gordon was blue
19:12
<Hixie>
man it's been too long
19:12
<Hixie>
(or possibly not long enough)
19:13
<Workshiva>
The former
19:13
<annevk5>
now you mention trains, apparently there's a Marklin shop here in Utrecht
19:13
<Workshiva>
Thomas is awesome
19:13
<annevk5>
thought of your train set when I saw that :)
19:14
<Hixie>
:-)
19:19
<jgraham>
gordon was indeed blue
19:19
<Workshiva>
Gordon was the fat Thomas
19:19
<Workshiva>
That's how I always thought of him
19:20
<Hixie>
thomas was a switcher, gordon was for long haul... though i don't think the people who wrote the stories understood the difference
19:21
<Lachy>
Percy was the green one.
19:22
gsnedders
finally realizes what you're on about
19:22
<gsnedders>
Oh man… Bunch of kids.
19:24
<Workshiva>
Yeah, that hurts bad coming from you
19:27
<jgraham>
Heh, I see that I didn't make abarth's list of people from browser vendors who are worth listening to
19:27
<jgraham>
I guess I should try harder or something
19:28
<Workshiva>
Maybe he doesn't know you're from a browser vendor
19:28
<jgraham>
I suppose that is possible
19:29
<jgraham>
But it seems unlikely
19:31
<Philip`>
jgraham: The list was only a "for example", and looks like it's intentionally listing one person per browser vendor
19:32
<gsnedders>
Hixie: How does you writing a separate Web Sockets spec to the IETF one help? Would you keep writing your spec if browser buy-in stuck with the IETF branch?
19:33
<Hixie>
gsnedders: no, if browsers aren't on board it's like with websql, i'd stop editing
19:33
<gsnedders>
Hixie: That wasn't quite clear on the list.
19:34
<Hixie>
well if i didn't i'd just be writing pointless fiction that didn't affect anyone anyway
19:34
<Hixie>
so it's rather moot
19:37
<jgraham>
Philip`: But I want to be important :p
19:38
<Workshiva>
jgraham: Clearly you need to eliminate the Opera employees before you in the ranking
19:38
<Workshiva>
That way you end up on the next list
19:39
<Philip`>
jgraham: You could be just a smidgen behind annevk5 on the perceived importance scale
19:39
<Philip`>
Or you could be right at the bottom
19:40
<Philip`>
so you should eliminate every single Opera employee, just to be sure
19:40
<jgraham>
Hmm
19:41
<jgraham>
How to kill the dutch?
19:41
<Hixie>
you'd be pretty important if you went on a rampaging homocide stream, but i'd urge you to consider if that's the right kind of importance for you
19:41
<jgraham>
Maybe make a series of tiny holes in their levees
19:41
<Hixie>
streak, even
19:41
<Workshiva>
You could crash the tulip market
19:45
<jgraham>
All of this sounds like too much effort really
19:45
<annevk5>
clearly abarth should be arrested for inciting violence
19:46
<jgraham>
I think I will just have to develop a zen-like perspective on my own insignificance
19:48
<annevk5>
is this my cue for saying you're not? or something? ;p
19:48
<jgraham>
No
19:48
<jgraham>
I am developing an indifference to it, remember
19:48
<jgraham>
If you say I'm not it will only confuse and upset me
19:49
<jgraham>
So I might go back to plotting to kill the Dutch
19:49
<annevk5>
I think I stand by my original statement
19:58
<jgraham>
I had forgotten about Percy
19:59
<jgraham>
Is it me or does Sordor sound uncomfortably like Mordor
20:00
<jgraham>
I would never have liked Thomas The Tank Engine so much if I had thought he was mainly carrying Orcs
20:00
<jgraham>
s/Sordor/Sodor/
20:00
<jgraham>
Which I guess makes a difference
20:01
<jgraham>
But still
20:21
<gsnedders>
Time to make me hate zcorpan again, and bring PHP html5lib up to date
20:24
<gsnedders>
Uh, the Python tests don't run for me
20:25
<gsnedders>
jgraham: You broke running tests with UTF-16 Python
20:38
<jgraham>
gsnedders: Ah, I think I expected that
20:38
<jgraham>
I may even have mentioned it in the commit log
20:38
<jgraham>
But I had no easy way to test
20:39
<jgraham>
gsnedders: (that is a lame excuse, yes, but I didn't really have time to fix it then)
20:40
<jgraham>
gsnedders: http://code.google.com/p/html5lib/source/detail?r=964568c175092c45156fe5a32a211e0d5d3781d8
20:40
<jgraham>
Probably
20:40
<gsnedders>
jgraham: There's lots of breakage, not just that
20:41
<jgraham>
Oh
20:41
<gsnedders>
Like, creating http://code.google.com/p/html5lib/source/detail?r=964568c175092c45156fe5a32a211e0d5d3781d8
20:41
<gsnedders>
Um, wrong clipboard
20:41
<gsnedders>
encode_entity_map
20:41
<gsnedders>
That throws an exception. :)
20:42
<gsnedders>
Which means import html5lib fails :)
20:42
<jgraham>
gsnedders: That has nothing to do with me
20:42
<jgraham>
Possibly
20:42
<jgraham>
Unless it was adding more entities that broke it
20:42
<jgraham>
Which is just silly
20:44
<gsnedders>
Adding non-BMP entities for the first time would
20:44
<gsnedders>
Now, to actually get tests passing instead of merely running
21:00
<gsnedders>
Huh, now I really don't get what's going on.
21:00
<gsnedders>
I appear to be hitting a data corruption bug in Python
21:02
<gsnedders>
Hah. This is awesome.
21:03
<gsnedders>
Negative lookbehind assertion in regexp causing data corruption.
21:04
<Philip`>
Got a test case?
21:05
<gsnedders>
Oh, no
21:05
<gsnedders>
I see what's going on
21:06
<gsnedders>
Hah, that is evil.
21:06
<gsnedders>
I can't write code.
21:07
<gsnedders>
Also: I just introduced a bug without breaking any tests.
21:07
<gsnedders>
We need more tests.
21:08
<jgraham>
What bug?
21:10
<gsnedders>
Stripping lone surrogate bytes would also strip the byte where the other half of the surrogate should be
21:10
<jgraham>
You would always remove two bytes rather than one?
21:12
<gsnedders>
Four bytes, two characters.
21:12
<jgraham>
So a test with {lone surrogate}{other} -> {replacemnt}{other} would be sufficient to test it
21:13
<gsnedders>
Yeah, I've added that
21:13
<gsnedders>
https://code.google.com/p/html5lib/source/detail?r=46df29539c714df260a18f280dfeaf96e7af62c5
21:13
<jgraham>
Hah, byte counting fail :)
21:15
<gsnedders>
I haven't tested on UCS4, but I've made no change to the code it uses effectively
21:15
gsnedders
wonders where his UCS4 build is
21:15
<jgraham>
gsnedders: I'm glad it was clear how I intended it to be patched for the UCS2 case :)
21:16
<jgraham>
gsnedders: Why remove the null from the regexp?
21:16
<jgraham>
Surely it is faster with it in?
21:17
<gsnedders>
More common code.
21:17
<jgraham>
Also, your patch is wrong
21:17
<gsnedders>
How?
21:17
<jgraham>
It doesn't take account of lone surrogates at the end of chunks
21:18
<gsnedders>
That's not a new issue
21:19
<jgraham>
Isn't it?
21:19
<gsnedders>
Well, we at least throw parse errors in that case
21:19
<gsnedders>
So it would fail tokenizer tests for that
21:20
<jgraham>
Can you actually end up with a non-lone surrogate at the end of the chunk in the UCS4 case?
21:21
<jgraham>
i.e. can you actually split the surrogate pair?
21:22
<jgraham>
It depends if we are reading bytes or characters
21:22
<gsnedders>
In the UCS4 case? No, you can never have a valid surrogate.
21:23
<gsnedders>
In the UCS2 case? Sure.
21:23
<jgraham>
Right, so we don't have the bug in the UCS4 case
21:24
<jgraham>
So the UCS2 patch is wrong in the sense that it misses a case that the UCS4 code covers
21:25
<gsnedders>
Indeed
21:25
<gsnedders>
But the UCS2 behaviour is already wrong in that case
21:25
<gsnedders>
And I've not made it any worse than it was before
21:25
<jgraham>
Agreed
21:25
<jgraham>
But the patch is still wrong :)
21:25
<gsnedders>
The patch is right, just incomplete. ;P
21:25
<jgraham>
However you want to think of it
21:26
<jgraham>
(it seems like since you are fixing it now, this would be a good time to make it right because otherwise we will have a subtle bug that will almost never happen but be reasonably surprising when it does)
21:28
<jgraham>
(you need to do roughly the same thing as the CR thing
21:29
<jgraham>
but bonus points for making it not that ugly)
21:30
<gsnedders>
It's harder than the CR thing
21:30
<gsnedders>
the CR thing is easy because you can just convert it to \n and ignore a LF in the next chunk
21:31
<gsnedders>
In this case you can't known what the right behaviour is until you get the next chunk… if there is a next chunk.
21:35
<Philip`>
Can you just stick the character onto the front of the next chunk?
21:35
<jgraham>
That might work
21:35
<Philip`>
or, uh, something like that
21:35
Philip`
has no idea how the code works really
21:35
<gsnedders>
Philip`: What if there's no next chunk?
21:36
<gsnedders>
(That's the problem with that solution)
21:36
<jgraham>
gsnedders: You just make sure that having a character in the unget buffer menas there is a next chunk
21:37
<jgraham>
Which I think is straightforward with the current code
21:40
<jgraham>
(just do data = self._danglingCharacter + self.dataStream.read(chunkSize)
21:41
<jgraham>
and then if some_regexp.match(data[-1]): self._danglingCharacter = data[-1]; data = data[:-1]; else: self._danglingCharacter = ""
21:41
<jgraham>
)
21:41
<jgraham>
or something
21:42
<jgraham>
There's not even any need for a regexp
21:43
<jgraham>
(just use ord)
21:44
<gsnedders>
jgraham: I don't get how to use the unget buffer for that
21:44
<gsnedders>
With how unget works, that is
21:46
<gsnedders>
Like, there is no buffer for unget
21:46
<jgraham>
gsnedders: I mean you have to create one
21:46
<jgraham>
that's self._danglingCharacter above
21:47
<jgraham>
Sorry, I don't think I was very clear
21:48
<gsnedders>
Really you're still not :)
21:49
<jgraham>
gsnedders: All I'm saying is
21:50
<jgraham>
If the last character is a \r or an unpaired surrogate, make a property that points to that character and slice it off the end of the chunk
21:50
<jgraham>
The next time we go to get a chunk, add it on the start
21:51
<jgraham>
this is sure to work because the signal for "we don't need no more chunks" is that readChunk returns nothing
21:51
<jgraham>
So there is a cost of one more cycle through readChunk if this is the last chunk
21:52
<jgraham>
but that isn't very common so it can be slow
21:52
<jgraham>
(we need to make sure we still do the right thing in that case of course)
21:53
<jgraham>
Then the only special magic we need is to make sure we detect when the last character is special and save it for next time
21:53
<jgraham>
Is that clearer, or am I talking nonsense?
21:58
<gsnedders>
That's clear
22:00
<jgraham>
I guess one special case is if the document is _only_ a \r character
22:00
<jgraham>
then you save the character but get an empty chunk back
22:01
<jgraham>
But you can probably deal with that where readChunk is called, or something
22:04
<jgraham>
Or in readchunk I guess
22:04
<jgraham>
Just by checking if length > 1 before you slice anything of
22:04
<jgraham>
f
22:04
<jgraham>
Which seems much simpler