| 11:42 | <annevk> | so only Mozilla has implemented ele.spellcheck and they do it as a boolean rather than enumerable attribute as the spec requires? |
| 11:42 | <jgraham> | annevk: Yes |
| 11:43 | <annevk> | I guess the spec wanted it to be in sync with .contentEditable |
| 11:43 | <jgraham> | hybi list fail - Thomas was blue not green. Henry was green. (and James was red) |
| 11:43 | <annevk> | maybe even on my request |
| 11:44 | <annevk> | jgraham, euh?! |
| 11:44 | <jgraham> | Yeah, I think the spec isn't going to happen |
| 11:45 | <jgraham> | annevk: Greg's extension of Hixie's transport analogy had a vehicle that is green, runs on rails, and answers to the name Thomas |
| 11:45 | <jgraham> | But Thomas was blue |
| 11:46 | <jgraham> | (possibly Thomas the Tank Engine is an Anglo-American curio) |
| 11:47 | <annevk> | Filed a bug on spellcheck. |
| 11:47 | <jgraham> | annevk: I already filed a bug indicating that throwing SYNTAX_ERR wasn't going to work |
| 12:25 | <MikeSmith> | annevk: I added an "HTML elements organized by function" section to the HtmlR doc - |
| 12:25 | <MikeSmith> | http://dev.w3.org/html5/markup/elements-by-function.html |
| 12:25 | <MikeSmith> | (I think you had suggested it should have one) |
| 13:04 | <annevk> | ah yeah |
| 13:04 | <erlehmann_> | annevk, I want to do some very light DOM manipulation in PHP and intend to use html5lib to get the DOM. Two question: First, is that a good choice? Second, what solution would you recommend to manipulate said DOM in PHP? |
| 13:04 | <annevk> | well, I suggested the main draft would be done in that way, but I suppose this works too |
| 13:04 | <annevk> | erlehmann_, I don't have experience with the PHP html5lib unfortunately |
| 13:05 | <annevk> | erlehmann, nor with PHP DOM manipulation :( |
| 13:05 | <annevk> | erlehmann, though overall that sounds like the best way if you plan on using PHP |
| 13:05 | <erlehmann> | just saw you as project owner on google code. that's the python part then, right? |
| 13:06 | <jgraham> | erlehmann: Well I think PHP isn't a good choice :) |
| 13:06 | <jgraham> | But if hat is a constraint then html5lib is a good choice for parsing the HTML |
| 13:06 | <jgraham> | as long as speed is not your main concern |
| 13:07 | <jgraham> | gsnedders and ezyang were mainly responsible for the PHP verson |
| 13:07 | <erlehmann> | jgraham, i agree wholeheartedly. when i applied for an internship in early 2008 and they asked me if i knew PHP, i told them why i hate it. |
| 13:07 | <annevk> | erlehmann, oh, yeah, I did the original Python version in part way back though jgraham knows and did more :) |
| 13:07 | <erlehmann> | but right now i have a gsoc project to finish. |
| 13:08 | <erlehmann> | and since i am writing a wordpress plugin … well ;) |
| 13:08 | <jgraham> | Ah |
| 13:08 | <jgraham> | You chose the wrong problem |
| 13:08 | <jgraham> | :) |
| 13:09 | <jgraham> | It is at least worth trying using PHP html5lib |
| 13:09 | <erlehmann> | probably. but all my friends are using wordpress. |
| 13:09 | <erlehmann> | and me too. though i will look into habari Really Soon Now [TM] |
| 13:10 | <jgraham> | That's PHP too, right? |
| 13:10 | <erlehmann> | yeah. i should probably bully my hoster into getting me some WSGI goodness so i can install a python-based imageboard instead of my boring old blog. |
| 13:13 | <erlehmann> | i'll look if PHP Simple HTML DOM Parser does it for me. i do not have that many edge cases and it looks nice and usable. |
| 14:17 | <gsnedders> | The PHP html5lib is really quite out of date |
| 14:17 | <gsnedders> | There's access to the libxml HTML parser from the DOM extension |
| 19:11 | <Hixie> | jgraham: yeah i considered saying gordon was green and thomas wasn't an eletric locomotive, but i figured that was maybe being pedantic about hte wrong thing :-) |
| 19:12 | <Hixie> | wait, gordon was blue |
| 19:12 | <Hixie> | man it's been too long |
| 19:12 | <Hixie> | (or possibly not long enough) |
| 19:13 | <Workshiva> | The former |
| 19:13 | <annevk5> | now you mention trains, apparently there's a Marklin shop here in Utrecht |
| 19:13 | <Workshiva> | Thomas is awesome |
| 19:13 | <annevk5> | thought of your train set when I saw that :) |
| 19:14 | <Hixie> | :-) |
| 19:19 | <jgraham> | gordon was indeed blue |
| 19:19 | <Workshiva> | Gordon was the fat Thomas |
| 19:19 | <Workshiva> | That's how I always thought of him |
| 19:20 | <Hixie> | thomas was a switcher, gordon was for long haul... though i don't think the people who wrote the stories understood the difference |
| 19:21 | <Lachy> | Percy was the green one. |
| 19:22 | gsnedders | finally realizes what you're on about |
| 19:22 | <gsnedders> | Oh man… Bunch of kids. |
| 19:24 | <Workshiva> | Yeah, that hurts bad coming from you |
| 19:27 | <jgraham> | Heh, I see that I didn't make abarth's list of people from browser vendors who are worth listening to |
| 19:27 | <jgraham> | I guess I should try harder or something |
| 19:28 | <Workshiva> | Maybe he doesn't know you're from a browser vendor |
| 19:28 | <jgraham> | I suppose that is possible |
| 19:29 | <jgraham> | But it seems unlikely |
| 19:31 | <Philip`> | jgraham: The list was only a "for example", and looks like it's intentionally listing one person per browser vendor |
| 19:32 | <gsnedders> | Hixie: How does you writing a separate Web Sockets spec to the IETF one help? Would you keep writing your spec if browser buy-in stuck with the IETF branch? |
| 19:33 | <Hixie> | gsnedders: no, if browsers aren't on board it's like with websql, i'd stop editing |
| 19:33 | <gsnedders> | Hixie: That wasn't quite clear on the list. |
| 19:34 | <Hixie> | well if i didn't i'd just be writing pointless fiction that didn't affect anyone anyway |
| 19:34 | <Hixie> | so it's rather moot |
| 19:37 | <jgraham> | Philip`: But I want to be important :p |
| 19:38 | <Workshiva> | jgraham: Clearly you need to eliminate the Opera employees before you in the ranking |
| 19:38 | <Workshiva> | That way you end up on the next list |
| 19:39 | <Philip`> | jgraham: You could be just a smidgen behind annevk5 on the perceived importance scale |
| 19:39 | <Philip`> | Or you could be right at the bottom |
| 19:40 | <Philip`> | so you should eliminate every single Opera employee, just to be sure |
| 19:40 | <jgraham> | Hmm |
| 19:41 | <jgraham> | How to kill the dutch? |
| 19:41 | <Hixie> | you'd be pretty important if you went on a rampaging homocide stream, but i'd urge you to consider if that's the right kind of importance for you |
| 19:41 | <jgraham> | Maybe make a series of tiny holes in their levees |
| 19:41 | <Hixie> | streak, even |
| 19:41 | <Workshiva> | You could crash the tulip market |
| 19:45 | <jgraham> | All of this sounds like too much effort really |
| 19:45 | <annevk5> | clearly abarth should be arrested for inciting violence |
| 19:46 | <jgraham> | I think I will just have to develop a zen-like perspective on my own insignificance |
| 19:48 | <annevk5> | is this my cue for saying you're not? or something? ;p |
| 19:48 | <jgraham> | No |
| 19:48 | <jgraham> | I am developing an indifference to it, remember |
| 19:48 | <jgraham> | If you say I'm not it will only confuse and upset me |
| 19:49 | <jgraham> | So I might go back to plotting to kill the Dutch |
| 19:49 | <annevk5> | I think I stand by my original statement |
| 19:58 | <jgraham> | I had forgotten about Percy |
| 19:59 | <jgraham> | Is it me or does Sordor sound uncomfortably like Mordor |
| 20:00 | <jgraham> | I would never have liked Thomas The Tank Engine so much if I had thought he was mainly carrying Orcs |
| 20:00 | <jgraham> | s/Sordor/Sodor/ |
| 20:00 | <jgraham> | Which I guess makes a difference |
| 20:01 | <jgraham> | But still |
| 20:21 | <gsnedders> | Time to make me hate zcorpan again, and bring PHP html5lib up to date |
| 20:24 | <gsnedders> | Uh, the Python tests don't run for me |
| 20:25 | <gsnedders> | jgraham: You broke running tests with UTF-16 Python |
| 20:38 | <jgraham> | gsnedders: Ah, I think I expected that |
| 20:38 | <jgraham> | I may even have mentioned it in the commit log |
| 20:38 | <jgraham> | But I had no easy way to test |
| 20:39 | <jgraham> | gsnedders: (that is a lame excuse, yes, but I didn't really have time to fix it then) |
| 20:40 | <jgraham> | gsnedders: http://code.google.com/p/html5lib/source/detail?r=964568c175092c45156fe5a32a211e0d5d3781d8 |
| 20:40 | <jgraham> | Probably |
| 20:40 | <gsnedders> | jgraham: There's lots of breakage, not just that |
| 20:41 | <jgraham> | Oh |
| 20:41 | <gsnedders> | Like, creating http://code.google.com/p/html5lib/source/detail?r=964568c175092c45156fe5a32a211e0d5d3781d8 |
| 20:41 | <gsnedders> | Um, wrong clipboard |
| 20:41 | <gsnedders> | encode_entity_map |
| 20:41 | <gsnedders> | That throws an exception. :) |
| 20:42 | <gsnedders> | Which means import html5lib fails :) |
| 20:42 | <jgraham> | gsnedders: That has nothing to do with me |
| 20:42 | <jgraham> | Possibly |
| 20:42 | <jgraham> | Unless it was adding more entities that broke it |
| 20:42 | <jgraham> | Which is just silly |
| 20:44 | <gsnedders> | Adding non-BMP entities for the first time would |
| 20:44 | <gsnedders> | Now, to actually get tests passing instead of merely running |
| 21:00 | <gsnedders> | Huh, now I really don't get what's going on. |
| 21:00 | <gsnedders> | I appear to be hitting a data corruption bug in Python |
| 21:02 | <gsnedders> | Hah. This is awesome. |
| 21:03 | <gsnedders> | Negative lookbehind assertion in regexp causing data corruption. |
| 21:04 | <Philip`> | Got a test case? |
| 21:05 | <gsnedders> | Oh, no |
| 21:05 | <gsnedders> | I see what's going on |
| 21:06 | <gsnedders> | Hah, that is evil. |
| 21:06 | <gsnedders> | I can't write code. |
| 21:07 | <gsnedders> | Also: I just introduced a bug without breaking any tests. |
| 21:07 | <gsnedders> | We need more tests. |
| 21:08 | <jgraham> | What bug? |
| 21:10 | <gsnedders> | Stripping lone surrogate bytes would also strip the byte where the other half of the surrogate should be |
| 21:10 | <jgraham> | You would always remove two bytes rather than one? |
| 21:12 | <gsnedders> | Four bytes, two characters. |
| 21:12 | <jgraham> | So a test with {lone surrogate}{other} -> {replacemnt}{other} would be sufficient to test it |
| 21:13 | <gsnedders> | Yeah, I've added that |
| 21:13 | <gsnedders> | https://code.google.com/p/html5lib/source/detail?r=46df29539c714df260a18f280dfeaf96e7af62c5 |
| 21:13 | <jgraham> | Hah, byte counting fail :) |
| 21:15 | <gsnedders> | I haven't tested on UCS4, but I've made no change to the code it uses effectively |
| 21:15 | gsnedders | wonders where his UCS4 build is |
| 21:15 | <jgraham> | gsnedders: I'm glad it was clear how I intended it to be patched for the UCS2 case :) |
| 21:16 | <jgraham> | gsnedders: Why remove the null from the regexp? |
| 21:16 | <jgraham> | Surely it is faster with it in? |
| 21:17 | <gsnedders> | More common code. |
| 21:17 | <jgraham> | Also, your patch is wrong |
| 21:17 | <gsnedders> | How? |
| 21:17 | <jgraham> | It doesn't take account of lone surrogates at the end of chunks |
| 21:18 | <gsnedders> | That's not a new issue |
| 21:19 | <jgraham> | Isn't it? |
| 21:19 | <gsnedders> | Well, we at least throw parse errors in that case |
| 21:19 | <gsnedders> | So it would fail tokenizer tests for that |
| 21:20 | <jgraham> | Can you actually end up with a non-lone surrogate at the end of the chunk in the UCS4 case? |
| 21:21 | <jgraham> | i.e. can you actually split the surrogate pair? |
| 21:22 | <jgraham> | It depends if we are reading bytes or characters |
| 21:22 | <gsnedders> | In the UCS4 case? No, you can never have a valid surrogate. |
| 21:23 | <gsnedders> | In the UCS2 case? Sure. |
| 21:23 | <jgraham> | Right, so we don't have the bug in the UCS4 case |
| 21:24 | <jgraham> | So the UCS2 patch is wrong in the sense that it misses a case that the UCS4 code covers |
| 21:25 | <gsnedders> | Indeed |
| 21:25 | <gsnedders> | But the UCS2 behaviour is already wrong in that case |
| 21:25 | <gsnedders> | And I've not made it any worse than it was before |
| 21:25 | <jgraham> | Agreed |
| 21:25 | <jgraham> | But the patch is still wrong :) |
| 21:25 | <gsnedders> | The patch is right, just incomplete. ;P |
| 21:25 | <jgraham> | However you want to think of it |
| 21:26 | <jgraham> | (it seems like since you are fixing it now, this would be a good time to make it right because otherwise we will have a subtle bug that will almost never happen but be reasonably surprising when it does) |
| 21:28 | <jgraham> | (you need to do roughly the same thing as the CR thing |
| 21:29 | <jgraham> | but bonus points for making it not that ugly) |
| 21:30 | <gsnedders> | It's harder than the CR thing |
| 21:30 | <gsnedders> | the CR thing is easy because you can just convert it to \n and ignore a LF in the next chunk |
| 21:31 | <gsnedders> | In this case you can't known what the right behaviour is until you get the next chunk… if there is a next chunk. |
| 21:35 | <Philip`> | Can you just stick the character onto the front of the next chunk? |
| 21:35 | <jgraham> | That might work |
| 21:35 | <Philip`> | or, uh, something like that |
| 21:35 | Philip` | has no idea how the code works really |
| 21:35 | <gsnedders> | Philip`: What if there's no next chunk? |
| 21:36 | <gsnedders> | (That's the problem with that solution) |
| 21:36 | <jgraham> | gsnedders: You just make sure that having a character in the unget buffer menas there is a next chunk |
| 21:37 | <jgraham> | Which I think is straightforward with the current code |
| 21:40 | <jgraham> | (just do data = self._danglingCharacter + self.dataStream.read(chunkSize) |
| 21:41 | <jgraham> | and then if some_regexp.match(data[-1]): self._danglingCharacter = data[-1]; data = data[:-1]; else: self._danglingCharacter = "" |
| 21:41 | <jgraham> | ) |
| 21:41 | <jgraham> | or something |
| 21:42 | <jgraham> | There's not even any need for a regexp |
| 21:43 | <jgraham> | (just use ord) |
| 21:44 | <gsnedders> | jgraham: I don't get how to use the unget buffer for that |
| 21:44 | <gsnedders> | With how unget works, that is |
| 21:46 | <gsnedders> | Like, there is no buffer for unget |
| 21:46 | <jgraham> | gsnedders: I mean you have to create one |
| 21:46 | <jgraham> | that's self._danglingCharacter above |
| 21:47 | <jgraham> | Sorry, I don't think I was very clear |
| 21:48 | <gsnedders> | Really you're still not :) |
| 21:49 | <jgraham> | gsnedders: All I'm saying is |
| 21:50 | <jgraham> | If the last character is a \r or an unpaired surrogate, make a property that points to that character and slice it off the end of the chunk |
| 21:50 | <jgraham> | The next time we go to get a chunk, add it on the start |
| 21:51 | <jgraham> | this is sure to work because the signal for "we don't need no more chunks" is that readChunk returns nothing |
| 21:51 | <jgraham> | So there is a cost of one more cycle through readChunk if this is the last chunk |
| 21:52 | <jgraham> | but that isn't very common so it can be slow |
| 21:52 | <jgraham> | (we need to make sure we still do the right thing in that case of course) |
| 21:53 | <jgraham> | Then the only special magic we need is to make sure we detect when the last character is special and save it for next time |
| 21:53 | <jgraham> | Is that clearer, or am I talking nonsense? |
| 21:58 | <gsnedders> | That's clear |
| 22:00 | <jgraham> | I guess one special case is if the document is _only_ a \r character |
| 22:00 | <jgraham> | then you save the character but get an empty chunk back |
| 22:01 | <jgraham> | But you can probably deal with that where readChunk is called, or something |
| 22:04 | <jgraham> | Or in readchunk I guess |
| 22:04 | <jgraham> | Just by checking if length > 1 before you slice anything of |
| 22:04 | <jgraham> | f |
| 22:04 | <jgraham> | Which seems much simpler |