00:08
<othermaciej_>
hober: PLH has some good factual points, but I think it's the first time I have seen the W3C object to even a WD-level spec being promoted...
00:10
<rubys>
I don't see the word "object" in PLH's post
00:12
<othermaciej>
most objections are not phrased in the form "I object"
00:13
<othermaciej>
let me put it this way, he spent more words on aspects he didn't like of Google's promotion of HTML5, than on aspects he liked
00:34
<Hixie>
hsivonen: yt?
00:40
<Hixie>
well i'll make this change and see if hsivonen complains, i guess
00:49
sayrer
is reminded of
00:49
<sayrer>
http://tomayko.com/writings/that-dilbert-cartoon
00:59
<Hixie>
othermaciej: your e-mail cuts off after "if Larry feels there is a resemblance between"
01:01
<othermaciej>
Hixie: thanks
01:12
Philip`
sees that Hixie is single-handedly making tens of percents of web pages less invalid than before
01:12
<Hixie>
i'm trying
01:57
<Hixie>
re http://www.w3.org/QA/2009/05/_watching_the_google_io.html -- maybe i should change the title on the whatwg version to "HTML5 Draft Standard" instead of "HTML5 Draft Recommendation" to remove any confusion... :-P
02:31
<theMadness>
Hixie, yes please.
02:31
<theMadness>
Change it HARD.
03:22
<roc>
hehe
03:24
<roc>
the latest Roy email is priceless
03:26
<ezyang>
linky?
03:28
<othermaciej>
roc: I find his "I'm the expert, stfu" attitude tiresom
03:29
<othermaciej>
roc: and I will reply to that effect on the list so I can be Shelley-compatible in my behavior
03:30
<roc>
yes
03:30
<roc>
plus the assertion about how browsers are only 1% of HTML clients
03:30
<ezyang>
gsnedders... you didn't unit test InputStream...
03:31
<ezyang>
This is an offense punishable... by death.
03:31
<othermaciej>
roc: I didn't even bother to argue with that
03:31
<othermaciej>
roc: but perhaps he is counting by distinct pieces of software as opposed to by number of users or volume of use
03:31
<roc>
yes, I think he is
03:32
<othermaciej>
which is not completely irrational, if he is using that to support an argument that the spec should be more general
03:32
<roc>
IIRC from threads long ago
03:32
<othermaciej>
but he still hasn't given any concrete examples
03:32
<othermaciej>
I'd *love* to hear how it's impossible for libwww-perl to comply with HTML5
03:32
<roc>
it's still irrational in at least two ways
03:33
<roc>
you've got to weight by how much an application is used, or you'll be swamped by never-used edge cases
03:33
<roc>
also, authors test in browsers, not any of these other hypothetical applications
03:33
<roc>
(maybe search engines, sort of)
03:34
<othermaciej>
as I see it, software that consumes HTML falls into one of two categories:
03:34
<othermaciej>
(a) it is intended to consume arbitrary HTML content from the Web
03:34
<roc>
so applications that want to process HTML need to process it the way authors do, if they want to capture the author's intent
03:34
<othermaciej>
(b) it is intended to consume only a restricted subset of the language or a specific set of documents in a controlled environment
03:35
<othermaciej>
it seems to me that category (a) must interoperate with browsers
03:35
<roc>
yeah
03:35
<othermaciej>
and for category (b), the spec is irrelevant
03:35
<roc>
right
03:35
<othermaciej>
if they want a relevant spec, they have to spec their subset
03:35
<othermaciej>
or they can just work ad-hoc without a spec
03:36
<othermaciej>
though I will add, that even for category (b) it is often desirable to use standard tools to create content and ordinary browsers to test
03:36
<othermaciej>
so even then, it may be highly desirable to interoperate with browsers
03:47
<ezyang>
Sweet. Extra five fails came from change in substr() behavior
03:47
<ezyang>
Now fixed
03:50
<ezyang>
Sigh, Google Code authorization is acting up again.
03:50
<ezyang>
Seriously dudes. You're google.
03:50
<ezyang>
You eat your own dogfood. (I hope)
03:51
<ezyang>
Strangely enough, the problem seems to fix itself when I navigate to my code.google.com/hosting/settings page :-/
03:59
<othermaciej>
Shelley quit again?
04:04
<ezyang>
Oh, awesome, the rest of the tokenizer fails are parse-error fails
04:08
<othermaciej>
roc: I guess if I were more inclined to be indignant I would be scandalized by Roy's claim that I am "clueless" about interpreters and "don't know anything about the field" of HTML
04:09
<roc>
you should be
04:09
<othermaciej>
I think it just makes him look like an idiot to say things like that
04:48
<olliej>
othermaciej: are you claiming you know stuff about interpreters again?
04:48
<ezyang>
Oh hey, whoops. Parse errors are not tokens
04:49
<othermaciej>
olliej: yeah, I like to pretend
04:50
<olliej>
othermaciej: :D
05:20
<ezyang>
Argh foster parenting
05:33
<mpilgrim>
hixie: html5/reddit alert: http://www.reddit.com/r/technology/comments/8nzok/google_waves_goodbye_to_email/c09w4u0
05:33
<mpilgrim>
that may overstate the situation somewhat
05:34
<mpilgrim>
would you settle for "html5 reference on reddit"?
05:38
<mpilgrim>
discussion of html5 video and codecs: http://www.reddit.com/r/programming/comments/8np9i/youtube_experimenting_with_html5_video/
05:38
<mpilgrim>
more video: http://www.reddit.com/r/programming/comments/8nr98/daily_motion_uses_html_5_video_with_ogg_theora/
05:40
<mpilgrim>
some meta-discussion: http://www.reddit.com/r/programming/comments/8ns2l/dear_proggit_html5_is_not_programming/
05:46
<othermaciej>
wow, lotsa discussion
05:46
<othermaciej>
(I believe Hixie has seen at least some of those threads)
05:47
<othermaciej>
mpilgrim: I
05:47
<othermaciej>
I'm waiting to see how your linking of those threads can be spun as malevolent
05:48
<jwalden>
1. knowledge is power. 2. power corrupts. 3. ... 4. GOOGLE PROFIT!
05:53
<othermaciej>
Google Wave looks a lot more visually busy than the typical Google web app
05:59
<roc>
I wonder why these codec discussions fail to mention that the moratorium on H.264 Internet "broadcasting" fees runs out next year
06:44
<othermaciej>
"HTML5 Could be the OS Killer"
06:44
<othermaciej>
apparently I'm suicidal
06:46
<othermaciej>
at the same time: "Google shows Native Client built into HTML 5"
06:51
ezyang
deals the death blow to the bug and emerges victorious
06:51
<ezyang>
Now, time for sleep
08:02
<Hixie>
wait... did shelley leave the wg again?
08:02
<Hixie>
man, she makes me dizzy
08:03
<othermaciej>
it seems that she is not presently a member of the HTML WG
08:04
<othermaciej>
according to the member list on the Web
08:06
<Hixie>
i guess it must be thursday
08:08
<othermaciej>
I assume she was offended by Sam pointing out that while she complains about outside-the-list snark/hostility, she also produces it
08:23
<Hixie>
othermaciej: no, i didn't intend the asymmetric case, i'll fix that in due course
08:24
<othermaciej>
Hixie: are you still going to allow UAs to switch between the two options at any time?
08:24
<Hixie>
othermaciej: yeah
08:24
<Hixie>
the requirement should be that for ports A and B
08:24
<othermaciej>
seems like a pretty low-value option if they would have to do so synchronously for both endpoints
08:24
<Hixie>
it doesn't have to be symmetric
08:24
<Hixie>
for port A, either A.owner->A or B->A
08:24
<othermaciej>
if you want to make it less of a loophole the requirement should be that A is either owned by B or by A's owner
08:24
<othermaciej>
right
08:25
<Hixie>
right now the text says either A.owner-> or A->B
08:25
<Hixie>
which is bogus
08:25
<Hixie>
the idea is still that you don't have to keep a reference to the object for it to survive long, since otherwise we expose GC
08:26
<Hixie>
"why do the first five messages get through?" "because after that it got around to GCing your port away" isn't a conversation i ever want to have :-)
08:26
<othermaciej>
so long as .close() gives you an out to clean up your resources, that's not so bad
08:26
<othermaciej>
but
08:26
<Hixie>
yeah
08:26
<othermaciej>
I don't see why exposing the potential unpredictability of GC is so bad
08:26
<othermaciej>
in an API to be used with concurrency
08:26
<othermaciej>
because concurrency means you can't really have deterministic guarantees of order of operations anyway
08:27
<Hixie>
you think this conversation is one that authors will be ok with? "why do the first five messages get through?" "because after that it got around to GCing your port away"
08:27
<othermaciej>
I wouldn't answer that way - I'd say "hold a reference to your MessagePort while you are using it"
08:27
<Hixie>
i think debugging such a bug would be a nightmare
08:27
<othermaciej>
debugging concurrency bugs is often a nightmare
08:28
<othermaciej>
imagine if your code subtly depends on whether worker A receives a message from worker B or window W first
08:28
<othermaciej>
you might never see the wrong order on "your" test system
08:28
<Hixie>
i agree that it has intrinsic pain
08:28
<Hixie>
adding more seems bad
08:28
<othermaciej>
this is why I think removing nondeterminism from an API for a concurrent system is wasted effort
08:29
<othermaciej>
well, by always either holding a reference or using .close() you can easily remove the risk
08:29
<othermaciej>
the only difference at this point is what happens to authors who are not careful in this way
08:30
<othermaciej>
do they get random memory leaks (result of the strong reference rule) or a bit of extra nondeterminism (result of not having such a rule at all)
08:30
<zcorpan_>
MikeSmith: "If the area element has no href attribute, then the area represented by the element cannot be selected, and the alt attribute must be omitted."
08:31
<othermaciej>
I don't think it's so important which happens, so I am not very concerned about the choice, as long as there is a reasonable way for content authors to do the right thing
08:32
<Hixie>
othermaciej: i think the memory leaks are the better option since if they are found to be common, browsers will just do the more complicated alternative.
08:32
<Hixie>
othermaciej: whereas they can't do anything about the other problem.
08:32
<othermaciej>
I don't think you understand how complicated the alternative is
08:32
<othermaciej>
if the leaks become a problem, I'd just violate the spec
08:34
<othermaciej>
(but hopefully giving authors the tools to DTRT will be sufficient)
08:35
<Hixie>
i don't think teh alternative is as bad as you make out. you just need to keep track of when port gets to zero js references, and when that happens, tell the other side the ID of the last message you received and the ID of the last message you sent, and say that you're ready to go away. when both sides do that and they both sent the same IDs, then they both know they can go away.
08:35
<Hixie>
consider the opposite problem -- one browser with 90% market share does GC every 20 seconds, the other does GC aggressively. a page that always sends two messages only always works on the former browser, always fails on the other.
08:36
<othermaciej>
JS isn't reference counted
08:36
<othermaciej>
there isn't a simple concept of "get to 0 references"
08:36
<othermaciej>
but there are also much more complex cases than you describe
08:36
<Hixie>
s/zero js references/would otherwise get sweeped/
08:36
<othermaciej>
consider a triangle of MessageChannels
08:36
<othermaciej>
where the MessagePorts in each window indirectly reference each other
08:37
<Hixie>
yeah, that's a tough one
08:37
<othermaciej>
(keep in mind also that linkage between three heaps is still a far simpler case for distributed GC than the fully general case)
08:39
<othermaciej>
(btw it's easy to get 3-way linkage like that through event listener closures capturing multiple message ports in scope)
08:39
<othermaciej>
(so it's not some bizzarre theoretical construct)
08:41
<Hixie>
agreed
08:42
<othermaciej>
so the upshot is that the opposite-endpoint ownership model is not really practical if you have separate heaps for workers, thus MessagePorts are immortal until you close() one of the endpoints or the other, or leave the Document
08:43
<othermaciej>
(er, for the last bit, I guess I should say, "until you navigate away")
08:44
<Hixie>
well exposing the GC just doesn't seem like an option to me
08:44
<Hixie>
i mean, we're going to extreme lengths to avoid it elsewhere
08:45
<othermaciej>
that just means that explicit resource management is the alternative
08:45
<othermaciej>
which is not necessarily that terrible
08:45
<Hixie>
yeah, i guess
08:45
<Hixie>
though people will screw that up too
08:45
<othermaciej>
just saying, it's not super helpful for the spec to make it seem like you can count on automatic resource management of these things
08:48
<Hixie>
hm, yes, we could make it more explicit that you should call .close() on ports
08:55
<Hixie>
http://beta.w3.org/standards/about.html
08:56
<Hixie>
after 15 years of telling people the w3c wasn't in the business of writing standards but was in fact in the business of writing recommendations, they're planning on confusing everyone by changing their mind
08:56
<Hixie>
good work w3c.
08:59
<annevk2>
seems like a positive change to me
09:01
<Philip`>
othermaciej: libwww-perl uses the HTML::HeadParser module which uses HTML::Parser and then looks for things like <meta http-equiv> and <base> and <link> and <isindex>
09:01
<othermaciej>
most people think of them as "Standards"
09:01
<Philip`>
(so they can be exposed via the HTTP header API)
09:01
<othermaciej>
calling them "Recommendations" and not standards was just confusing
09:01
<othermaciej>
Philip`: I did not know that
09:03
<othermaciej>
Philip`: is HTML::HeadParser // HTML::Parser unable to conform to HTML5?
09:03
<othermaciej>
(I assume not, since people have written HTML5 parsers in perl...)
09:04
<Philip`>
othermaciej: I see no reason that it couldn't use an HTML5 parser
09:04
<Philip`>
(though I don't know whether that's enough to technically "conform to HTML5")
09:05
<othermaciej>
I would guess it falls under the 'Data mining tools' conformance class
09:05
<othermaciej>
though perhaps there should be a conformance class for reusable parser libraries
09:05
<othermaciej>
since such a library is testable by itself usually, but can't predict what kind of software will use it
09:06
<Philip`>
Seems hard to specify since the libraries will all have different output formats
09:06
<othermaciej>
Philip`: anyway - thanks for pointing that out, though overall I still think Roy is wrong in his assertions, and incredibly rude
09:06
<othermaciej>
well, XML has rules for XML processors and people are able to determine whether XML parser libraries conform
09:07
<Philip`>
I guess it would have to specify equivalence between the conceptual DOM created by the spec's algorithm, and the concrete DOM/SAX/ElementTree/etc created by implementations
09:07
<Philip`>
or just leave it undefined and tell people to use common sense
09:10
<Philip`>
http://translate.google.com/translate?hl=en&sl=ru&tl=en&u=http://robotstxt.org.ru/rurobots/yandex
09:10
<Philip`>
"Yandex Robot supports tag noindex, which prohibits job Yandex index queries (office) of the text. At the beginning of a service fragment is <noindex>, but in the end - </ noindex>, Yandex, and will not index the site text."
09:11
<zcorpan_>
apache indexes would benefit from using an html5 parser here http://simon.html5.org/test/html/parsing/fragment/content-model-flag/
09:12
<zcorpan_>
(it shows just "setting innerHTML" instead of the correct "setting innerHTML on <title>")
09:12
<zcorpan_>
"setting innerHTML on"
09:13
<zcorpan_>
nowadays i try to remember escaping < as &lt; in titles to cater for the apache bug but it would be nice if it worked correctly
09:21
<zcorpan_>
Hixie: how was http://www.w3.org/Bugs/Public/show_bug.cgi?id=6586 fixed?
09:22
<Hixie>
it was fixed lng ago
09:22
<Hixie>
i just forgot to mark the bug as fixed
09:22
<Hixie>
spec now says:
09:22
<Hixie>
# When a Document is in quirks mode, margins on HTML elements at the top or bottom of the initial containing block, or the top of bottom of td or th elements, are expected to be collapsed to zero.
09:23
<Hixie>
othermaciej: your e-mail to roy had an incomplete sentence again
09:23
<Hixie>
"If they can only handle a defined subset inside their walled garden, then they are not conforming implementations. If they are"
09:23
<othermaciej>
Hixie: in case you coudn't tell, I edit out of order a lot
09:23
<zcorpan_>
Hixie: the spec quotes that exact text and says it's wrong
09:24
<zcorpan_>
er
09:24
<othermaciej>
*couldn't
09:24
<zcorpan_>
the bug
09:24
<Hixie>
othermaciej: so do i, but then i proof-read :-P
09:24
<othermaciej>
I do too - had a real bad headache today though
09:24
<othermaciej>
surprised I was able to be coherent at all
09:24
<othermaciej>
I should do a TL;DR reply to myself
09:24
<Hixie>
zcorpan_: no, the text in the bug is different
09:24
<othermaciej>
because in all the arguing, I do have an interesting main point
09:24
<Hixie>
othermaciej: ah, sorry to hear that
09:24
<Hixie>
about the headache, not the point
09:25
<othermaciej>
which is that HTML processors either (a) are expected to work with general public Web content, in which case they must interoperate with browsers, or (b) work with a restricted subset of content (either conforming to some subset rule or in a walled garden) in which case the spec is irrelevant to them
09:27
<zcorpan_>
Hixie: ok it's different, but i don't see how it collapses when the body has a border
09:28
<Hixie>
zcorpan_: why would the border affect it?
09:28
<Hixie>
the margins just collapse to zero whatever else is there
09:28
<Hixie>
with whatever else, i mean
09:29
<zcorpan_>
Hixie: the body is not the initial containing block, and margin collapsing doesn't happen when there's a border, so as i read it it wouldn't collapse when there's a border
09:30
<zcorpan_>
Hixie: instead the body's margins would collapse, which they shouldn't
09:30
<Hixie>
oh it should say top or bottom of the <body>, not the ICB?
09:30
<zcorpan_>
yeah
09:33
<Mrmil>
Dear WHATWG, I'm yet-another-I-want-to-help person and I would like to ask you kindly for your feedback.
09:33
<Mrmil>
I'm working on a case-study, a common web presentation homepage written in HTML 5, which (when finished) can get chopped up into smaller parts and form a tutorial.
09:33
<Mrmil>
Please note it's not finished at all, many things need to be done and I need to educate myself, too. :) If you are interested, you can find it here: http://server.ebrana.cz/olda/_apps/html5/ .
09:33
<Mrmil>
It's a temporary URL until I/we decide what to do with it next.
09:33
<Mrmil>
I'm planning on adding more feature to this page, so I'll post here again when done. If you want to provide feedback, please email me at vetesnik⊙mc And yeh, I'll buy you a cup of coffee. :) Thanks.
09:34
<zcorpan_>
mmm coffee
09:35
<zcorpan_>
Mrmil: http://gsnedders.html5.org/outliner/
09:36
<Mrmil>
zcorpan_: Ooo, cool, thanks!
09:36
<zcorpan_>
Mrmil: <html lang="cs"> - shouldn't it be "en"?
09:37
<Mrmil>
zcorpan_: Yes, I'm czech so I forgot to change it.
09:37
<Mrmil>
zcorpan_: done
09:37
<hsivonen>
Philip`: I thought the correspondence of DOM/SAX/ElementTree etc. to infoset is already understood
09:37
<zcorpan_>
<section id="header"> - without checking the outline or further in the source, knee-jerk reaction is that this should be <header>
09:38
<hsivonen>
Philip`: and HTML5 defines how to coerce the output of the parsing algorithm to infoste
09:39
<zcorpan_>
hsivonen: still people say "HTML5 talks about DOM, there are tools that don't use a DOM, hence HTML5 is useless for those tools"
09:39
<hsivonen>
zcorpan_: It's a sign of people not having read the spec carefully
09:40
<zcorpan_>
Mrmil: the <h1>Navigation</h1> should probably go inside the <nav>
09:40
<jgraham>
Mrmil: <p class="skipLinks"> should be unnecessary (the whole paragraph)
09:41
<zcorpan_>
<div id="wrapper"> is unnecessary too (style <body> instead)
09:41
<jgraham>
<section id="content"> -> <article> maybe
09:41
<jgraham>
<section id="sidebar"> -> <aside> aiui
09:41
<Philip`>
hsivonen: Coercion to infosets doesn't seem an ideal way to define equivalence, since the coercion can lose information
09:42
<jgraham>
Philip`: I'm still not sure I understand what information is lost
09:42
<zcorpan_>
Mrmil: i would probably have a single <nav> and have subheadings for quicknav and language
09:42
<Philip`>
jgraham: It can e.g. drop attributes whose name starts with "xmlns"
09:42
<zcorpan_>
s/subheadings/subsections/
09:42
<jgraham>
Philip`: Oh yeah, I had forgotten that.
09:43
<hsivonen>
Philip`: I think common sense should be enough to fill in the gaps given the stream to DOM spec and the coercion spec, but I guess it could be more explicit
09:43
<zcorpan_>
Mrmil: <hr class="displayNone"> - clearly presentational abuse of class
09:43
<hsivonen>
Philip`: is there any implementor who hasn't been able to figure out how to map the non-infoset parts of the DOM to an arbitrary non-infoset-enforcing API?
09:43
<Philip`>
Maybe it should be explicit that you should use common sense
09:44
<hsivonen>
common sense is rare, though :-(
09:44
<Philip`>
e.g. add a conformance class for "HTML parser library" and require that it has a common-sense equivalence to the DOM produced by the specified algorithm
09:44
<jgraham>
zcorpan_: Presenational "abuse" of class isn't forbidden is it?
09:44
<zcorpan_>
How to read this specification: use common sense
09:44
<jgraham>
(but I think that <hr> should go)
09:44
<Mrmil>
zcorpan_: I was wondering about that, I'll switch it to the <nav>
09:45
<hsivonen>
Philip`: common-sense equivalance or where prohibited, using the coercion section's rules
09:45
<zcorpan_>
jgraham: no it's not forbidden but it's not good practice
09:45
<Philip`>
hsivonen: That sounds sensible
09:45
<Mrmil>
jgraham: what would you suggest then?
09:45
<hsivonen>
Philip`: I think I could formulate general rules for any API using the infoset as an aid
09:45
<Mrmil>
jgraham: removing whole skiplinks or removing the <p>?
09:45
<jgraham>
Mrmil: Why do you need an <hr>?
09:45
<hsivonen>
because the problems are mostly arbitrary restrictions on what characters are allowed where
09:46
<zcorpan_>
Mrmil: the arvhices list in the sidebar probably doesn't need <nav>
09:47
<hsivonen>
Philip`: 1) figure out how the infoset maps to API X
09:47
<zcorpan_>
Mrmil: <section id="footer"> - <footer>
09:47
<Philip`>
I suppose I just think it'd be nice to be able to say "this is a conforming HTML5 parser library according to the spec", instead of having to say "this is an HTML5 parser library which could be used as part of a data mining tool and would not by itself cause the data mining tool to be non-conforming"
09:47
<hsivonen>
Philip`: 2) Apply the DOM to infoset coercion from DOM Level 3
09:47
<Hixie>
nn
09:47
<jgraham>
Mrmil: Remove the skiplink. AIUI they don't work that well in practice an in theory the browser can guess rather well anyway (e.g. look for the first <article>)
09:47
<hsivonen>
Philip`: in step #2, pretend that restrictions on what characters are allowed where don't apply
09:48
<Mrmil>
jgraham: hr's removed, OK, will remove skiplinks too :)
09:48
<jgraham>
Mrmil: <section class="todo"> no need for the wrapper <div>
09:48
<Mrmil>
jgraham: done
09:48
<hsivonen>
Philip`: 3) Map infoset to API X per step 1
09:48
<Mrmil>
jgraham: I use the div only for styling purpose
09:48
<zcorpan_>
Mrmil: "© Company Name 2009, All rights reserved. Company Product" doesn't make sense as a paragraph
09:48
<hsivonen>
Philip`: 4) If this causes API X to fail, remove failures by applying the coercion section
09:49
<hsivonen>
Philip`: is that comprehensive enough?
09:50
<jgraham>
Mrmil: I bet you don't really need it
09:50
<hsivonen>
Hixie: you should make a reference to DOM Level 3 for the baseline DOM to infoset mapping that the coercion rules modify when needed
09:50
<Philip`>
hsivonen: I guess that could work
09:51
<hsivonen>
Philip`: works for everything but RDFa
09:51
<Mrmil>
jgraham: you bet right, but what if I got a design with rounded corners?
09:51
zcorpan_
ponders about <address> in the sidebar
09:51
<jgraham>
Mrmil: border-radius?
09:52
<Mrmil>
jgraham: doesn't work in ie, not implemented enough, cannot use it for production use, my superiors would kill me
09:52
<Mrmil>
jgraham: will have to wait for css3 a little while
09:53
<jgraham>
Mrmil: I thought this was a case study not a production site for clients who demand precidely the same look in all browsers
09:53
<jgraham>
*precisely
09:53
<Mrmil>
jgraham: Well it's a case study based on real demands, I don't want to make it a SCI-FI, people need real tutorials for real needs, I should've said that I guess
09:54
<zcorpan_>
Mrmil: the link farm at the bottom should probably go in the <footer>, too
09:56
<Mrmil>
zcorpan_: Are all the things you said alright now?
09:56
<hsivonen>
zcorpan_: are you planning on having a mapping from Web DOM to Infoset?
09:57
<Mrmil>
zcorpan_: "? Company Name 2009, All rights reserved. Company Product" doesn't make sense as a paragraph -> what would you suggest then? I don't like the idea of <div>ing everthing
09:58
<hsivonen>
Mrmil: it's a "paragraph" in the HTML5 sense of "paragraph"
09:58
<zcorpan_>
Mrmil: <p><small>copyright</small></p><p><a><img>...
10:00
<zcorpan_>
hsivonen: dunno
10:01
<zcorpan_>
hsivonen: i guess it would be good
10:01
<Mrmil>
zcorpan_: better now? I'm a bit confused :)
10:01
<zcorpan_>
Mrmil: yes, now at least the copyright paragraph makes sense. :)
10:02
<Mrmil>
zcorpan_: ok :)
10:03
<zcorpan_>
now where's my coffee?
10:03
<Mrmil>
zcorpan_: let me know if there are any other problems, I'll have some lunch now. Thanks for the feedback. Mmm, how do you like your coffee?
10:03
<zcorpan_>
Mrmil: black
10:04
<Mrmil>
zcorpan_: there you go http://www.nothinggeek.com/wp-content/uploads/2009/01/black-coffee.jpg
10:04
<zcorpan_>
thanks!
10:04
<Mrmil>
Cheers. I'll have some after lunch. Love it.
10:05
<Mrmil>
I have a question: should there be only one <nav> on a page?
10:05
<zcorpan_>
there's no such rule
10:05
<zcorpan_>
but i guess it makes the page easier to navigate
10:06
<Philip`>
zcorpan_: Maybe the TAG could give you some advice on the difference between coffee and a representation of coffee
10:07
<Philip`>
(Seems a similar problem to http://www.w3.org/2001/tag/issues#httpRange-14)
10:07
<Mrmil>
zcorpan_: On some of our presentations, we have a header-menu which goes throughout whole site, and then we have a column menu which goes into subpages for the particular page. Does that mean that both of them should be in a nav? I guess so.
10:11
<Mrmil>
jgraham: <section id="content"> -> <article> -> I was thinking about that. Then if I have real news up there, I'll nest <article> right?
10:12
<Mrmil>
jgraham: <section id="sidebar"> -> <aside> aiui -> I was thinking about that too, but wasn't sure
10:12
<jgraham>
Mrmil: Well it should only be <article> if it really is an article
10:12
<jgraham>
If it is a collection of articles <section> is probably fine
10:13
<Mrmil>
jgraham: it's not always an article. The contents of the #content part varies - depending whether you are on homepage, detail product, checkout, etc.
10:13
<jgraham>
Mrmil: If it is a collection of things then <section> probably works better
10:13
<Mrmil>
jgraham: Ok.
10:13
<jgraham>
(a collection on the same page)
10:15
<Mrmil>
jgraham: is my #sidebar had a product catalogue, bestselling products or banners, should it still be an <aside>?
10:15
<jgraham>
Like <section><h1>Blog posts</h1><article><h1>My first Blog Post</h1></article><article><h1>My second blog post</h1></article>
10:15
<jgraham>
</section>
10:16
<jgraham>
Is better than it would be with s/section/article/
10:16
<jgraham>
But <article><h1>Product details</h1></article> is fine
10:17
<jgraham>
Mrmil: Per spec, yes
10:17
<jgraham>
(re: <aside>)
10:18
<annevk2>
Hixie, reason for quitting seems other work: http://twitter.com/burningbird/status/1953228067
10:19
<zcorpan_>
Mrmil: i usually have a sub list for that case i.e. <nav><ul><li><a>home</a><li><a>products</a><ul><li>product 1<li><a>product 2</a></ul><li><a>etc
10:22
<zcorpan_>
Philip`: i can always print the representation of coffe to obtain a physical form
10:23
<zcorpan_>
coffee
10:36
<annevk2>
"Hurray for the tracking view!" nice to know it's appreciated :)
10:37
<annevk2>
it's unfortunate that making it better means a significant increase in complexity (afaict)
10:40
<jgraham>
annevk2: Where are you quoting?
10:42
<Philip`>
jgraham: http://lists.whatwg.org/htdig.cgi/whatwg-whatwg.org/2009-May/019982.html
11:01
<Mrmil>
Ok, any other suggestions for http://server.ebrana.cz/olda/_apps/html5/ ?
11:03
<jgraham>
<section id="header"> -> <header>
11:04
<jgraham>
You could probably remove <section id="content"> altogether
11:04
<jgraham>
(just keep the articles)
11:04
<jgraham>
Since you don't have another <h1> that applies to the <body>
11:05
<jgraham>
The empty <span> elements are kind of ugly
11:06
<jgraham>
Having both <section id="footer"> and <footer> seems odd
11:11
<Mrmil>
jgraham: removing <section id="content"> might be a problem due to 2-col layout and 3-col layout problems, but will try it occasionally.
11:12
<Mrmil>
jgraham: empty span elements are ugly but they convey meaning so I have to keep the text behind it so it's accessible with images disabled. And I don't want to use <img>'s for that, that's even more ugly. :)
11:15
<Mrmil>
jgraham: I changed section id="header" to header id="header", I need the id so I don't style every header this way. What would you suggest do with the footers? I'd change the section id="footer" into footer element but then there would be 2 footer's. The link farm is ugly but our IM dept. adds it everywhere so I have to deal with it somehow.
11:19
<annevk2>
http://twitter.com/gazcoop/statuses/1957781728 :)
11:21
<annevk2>
http://twitter.com/martin_probst/statuses/1958288180 -- "HTML5 Spec is a weird document. Offline Applications section is almost just C code, nothing about intent, expected outcomes or behaviour." hard to argue with that; hopefully the introduction section addresses these concerns adequately
11:36
<Mrmil>
jgraham: One more qustion, if I want to keep the #content section just for styling purposes then, should I make it <div> instead?
11:36
<hsivonen>
Whew. The check-in message for r3148 is wrong.
11:36
<hsivonen>
Spec looking good.
11:42
<annevk2>
Google Wave looks pretty neat
11:42
<annevk2>
little bit scared of all the probable data lock-in though
11:43
<zcorpan>
http://www.techcrunch.com/2009/05/28/google-wave-the-full-video-from-google-io/ - drag and drop several images from desktop to web app in one go
11:44
<jgraham>
annevk2: Isn't it supposed to be an open protocol somehow?
11:45
<annevk2>
I was not assuming open protocol means that really
11:45
<annevk2>
but I don't know the details of that yet
11:48
<jgraham>
Mrmil: Re: #content it probably doesn't matter much, but per spec, if the <body> doesn't have a <h1> then it will look like an empty section in an outline
11:50
<jgraham>
But I'm trying to get Hixie to change that
11:51
<jgraham>
(because I expect a bunch of people to do essentially the same thing that you have done)
12:05
<gsnedders>
ezyang: But all the Tokenizer test cases test it :P
12:06
<annevk2>
I guess since they open source it it's ok
12:07
<annevk2>
Hopefully the network effects are still good enough if you run your own...
12:07
<gsnedders>
""\"\"\"\"geoffers\"\"\"@gmail com — that's certainly an interesting email address to receive from
12:08
<Philip`>
annevk2: http://www.waveprotocol.org/ - sounds like the idea is that you can be part of the global network without being locked into Google's servers
12:09
<annevk2>
yeah
12:09
<Philip`>
Seems like it's basically XMPP (Jabber) with some extensions
12:09
<annevk2>
TAG might enjoy speculating about that versioning issue :)
12:09
<Philip`>
so it should have the same level of openness as XMPP, with everyone able to link to each other's servers
12:11
<Philip`>
and the extensions seem to be a way of encoding deltas for XML documents
12:11
<Philip`>
(in a composable way)
13:00
<annevk2>
I wonder how HTML WG discussions would look in Wave. I have the feeling it might be pretty overwhelming.
13:04
<jgraham>
annevk2: I was thinking the same thing
13:05
<Philip`>
HTML WG discussions are pretty overwhelming regardless of the medium
13:06
<annevk2>
I wonder how URL addressing works with Wave. I guess you'd have a "bot" that is invited in Wave discussions that pushes content out now and then...
13:26
<Lachy>
wow, I just noticed that pave the cowpaths is getting discussed again. Haven't read the whole thread yet, guessing it's going to be a lot of nonsense and misunderstanding again
13:55
<gsnedders>
"How do you work this [computer]? You go click! Oh fuck it, why doesn't it work, I thought if you put the thing [cursor] there [over text field] it would go there!"
13:58
gsnedders
sighs
13:58
<gsnedders>
being around my mother using a computer is stressful.
13:59
<hsivonen>
I want to complain about HotSpot's 8000-byte limit, but you've heard it already
13:59
<hsivonen>
seems crazy to have to tweak code around a magic number like that
14:02
<jgraham>
hsivonen: I promise to say nothing about proposals about charset declerations only working in the first 512/1024/some other fixed number of bytes
14:03
<jgraham>
;)
14:03
<hsivonen>
jgraham: that's totally different!
14:04
<Philip`>
Is there no command-line option to change HotSpot's behaviour?
14:05
<hsivonen>
Philip`: -XX:-DontCompileHugeMethods
14:06
<hsivonen>
Philip`: but requiring library users to know about a flag like that would not be cool
14:06
<hsivonen>
I wonder if AppEngine has a limitation like that
14:07
hsivonen
starts refactoring the tokenizer loop for better perf, line numbers in C++ and better localizability
14:07
<gsnedders>
http://codingforums.com/showthread.php?t=167485 — anyone help?
14:08
<hsivonen>
oh, and CRLF and \0 suck too
14:11
<hsivonen>
maybe I can start an HTML optimization meme about how much faster HTML parser if it has no CR characters and has LFs instead
14:12
<hsivonen>
s/parser/parses/
14:12
<jgraham>
gsnedders: Don't know but you are describing a partial solution rather than a problem. If you describe a problem you may get better help
14:12
<gsnedders>
jgraham: The problem is calling a function on click except when the click falls on an element with a default action!
14:13
<jgraham>
gsnedders: I doubt it. That's part of a solution to some other problem
14:13
<gsnedders>
jgraham: I want to be able to click on a block to hide and show content.
14:38
<hsivonen>
w00t! refactoring error messages saves hundreds of byte codes
14:38
<hsivonen>
(innocent string appends compile into a lot of byte codes)
14:39
<ezyang>
gsnedders: Nonsense!
14:39
<gsnedders>
ezyang: Yes, I'm a bad little boy.
14:39
<ezyang>
gsnedders: More seriously, if there ever is a failing test-case in Tokenizer, it's nice to know that it actually is Tokenizer's fault, and not InputStream's. A test-suite goes a way to do this
14:39
<gsnedders>
ezyang: Indeed
14:40
<gsnedders>
ezyang: Some of it was really not very fun to debug while splitting it out with obscure bugs in the input stream
14:40
<ezyang>
Yeah. Kudos on getting Parse Errors to mostly work
14:41
<ezyang>
Like, that's some pretty in-depth work you did
14:41
<ezyang>
(also, I'm ignoring parse errors completely for TreeConstructer, so you get to do it again :-)
14:41
<gsnedders>
ezyang: But I couldn't decide how to do a test suite for something where we weren't using JSON, etc. for input. We'd probably end up disagreeing how to do it. :P
14:42
<gsnedders>
ezyang: Have you merged back into trunk yet?
14:42
<ezyang>
Yeah, it's all in the trunk
14:42
gsnedders
ought to do revision for his computing exam
14:42
<ezyang>
Re test suite: I dunno; I think that the class is simple enough that pure PHP works
14:42
<ezyang>
I thought you finished your exams?
14:42
<gsnedders>
Actually, learning for the first time some of it.
14:43
<gsnedders>
No, computing next Thursday.
14:50
<hsivonen>
form feeds are almost as bad as CRs
14:57
<hsivonen>
http://www.builderau.com.au/news/soa/Google-Chrome-gets-HTML-video-support/0,339028227,339296704,00.htm
15:01
<Lachy>
hsivonen, I wouldn't object to the validator issuing warnings about the use of CR or CRLF.
15:01
<Lachy>
it's unfortunate, though, that nothing you do could ever lead to CR being abolished entirely :-(
15:02
<jgraham>
Lachy: Windows users wouldn't like that
15:03
<Lachy>
jgraham, only Notepad users would have serious problems with it. But they have bigger issues to worry about anyway.
15:03
<hsivonen>
Lachy: yeah, it's unfortunate :-(
15:04
<hsivonen>
to make speculative parsing work sanely, the only solution I've come up with makes CR slower than LF
15:04
<jgraham>
Also, "Consider use cases" doesn't convey the spirit of "pave the cowpaths" at all
15:04
<Lachy>
I wonder what it would take to get Microsoft to start migrating to the use of LF only
15:04
<Lachy>
jgraham, yes it does, since pave the cowpaths is entirely about use cases
15:05
<Lachy>
well, at least, it's meant to be
15:05
<jgraham>
Lachy: No it's about solutions
15:05
<Lachy>
no it's not!
15:05
<jgraham>
It is.
15:05
<hsivonen>
hmm. I think I'm not going to support speculative parsing and coercion into XML infoset in the same parser instance
15:06
<jgraham>
Lachy: There should probably be a principle like "Work from use cases" but it doesn't mean anything like "pave the cowpaths"
15:06
<othermaciej>
I agree with jgraham
15:06
<othermaciej>
"Work from use cases" should be a principle
15:07
<othermaciej>
"solve real problems" is meant to convey that idea but its title is needlessly confrontational
15:07
<Lachy>
jgraham, it's about looking at what authors are trying to do and providing solutions, possibly based on the existing solution, or providing a better solution
15:07
<othermaciej>
"pave the cowpaths" is all about "if authors do X a lot in markup, then it's probably a good idea to enable that, or something close to it, as a feature"
15:07
<Lachy>
othermaciej, that's not what it's meant to be.
15:07
<othermaciej>
"as opposed to making up a wildly different solution to the same thing"
15:08
<othermaciej>
well, I put it in the Design Principles document, back when it was a wiki page
15:08
<jgraham>
Lachy: That is part of it but "pave the cowpaths" implies that common practice should be given more consideration than it would if it was not in common practice
15:08
<othermaciej>
I would like to think I knew what I was saying
15:08
<othermaciej>
it's inspired by the microformats principle of the same name
15:08
<jgraham>
e.g. we would never invent <br/> if it wasn't already common practice but since authors want to do it, we allow it
15:09
<Lachy>
yeah, and I remember tantek coming in here once and insisting that cowpaths was about use cases, and that that's how it's applied to microformats
15:10
<othermaciej>
microformats wiki says:
15:10
<othermaciej>
"Remember, we're paving the cowpaths- before you do that you have to find the cowpaths. Your examples should be a collection of real world sites and pages which are publishing the kind of data you wish to structure with a microformat. From those pages and sites, you should extract markup examples and the schemas implied therein, and provide analysis."
15:10
<jgraham>
Lachy: It is about use cases in the sense that a cow path must be something that people want to do
15:10
<jgraham>
and already do do
15:10
<Lachy>
http://krijnhoetmer.nl/irc-logs/whatwg/20070812#l-86
15:10
<othermaciej>
HTML5 also considers use cases that aren't anything anyone does
15:10
<Lachy>
jgraham, yes
15:10
<othermaciej>
those would not be a "cow path"
15:11
<annevk2>
http://microformats.org/wiki/process#Document_Current_Behavior
15:11
<othermaciej>
(or rather, aren't anything anyone does, because you can't yet)
15:11
<jgraham>
Lachy: But it explicitly favours adopting solutions that are close to the de-facto solutions even if they are not the ones that you would design with a clean-slate approach
15:12
<othermaciej>
anyway, I would be happy to remove the exact phrase "pave the cowpaths" if only so I never have to hear accessibility enthusiasts argue that some markup feature is or isn't a cowpath
15:12
<Lachy>
jgraham, to the extent that the existing solutions are not significantly problematic, yes.
15:13
<othermaciej>
but I think the idea should be captured
15:13
<othermaciej>
and it's not the same as "work from use cases"
15:13
<jgraham>
Lachy: Of course. "Favours" is not absolute. It is, as always, a trade off
15:14
<jgraham>
othermaciej: I agree that e should remove the phrase "pave the cowpaths". Whilst it seems intuitive to me what it implied by such a principle it seems like it creates a great deal of confusion
15:14
<jgraham>
Maybe it should be captured by a principle like "Support common practice" or something
15:15
<Lachy>
How about "Analyse Existing Practices"
15:15
<Lachy>
or what jgraham suggested
15:15
<jgraham>
I was just about to suggest "investigate existing practice"
15:16
<jgraham>
:)
15:16
<Lachy>
either of those work
15:18
<jgraham>
Maybe with the principle saying something like "[...] where the existing solution does not have significant drawbacks consider adopting it rather than forbidding it or creating something new"
15:19
<Lachy>
that's close to Don't Reinvent the Wheel
15:20
<jgraham>
Lachy: You would need more text at the start to explain that you have to investigate the ways that things are already being done
15:21
<jgraham>
(my main point was that an explicit disclaimer that solutions with big problems should not be adopted wholesale)
15:21
<Lachy>
ok
15:29
<Lachy>
Something like this might work as part of the description "Investigate existing, de-facto solutions to problems and evaluate them in regards to the use cases being addressed. Consider either adopting or developing solutions based on the existing practices."
15:30
<Lachy>
(in addition to something like what jgraham suggested)
15:33
<annevk2>
Don't Reinvent the Wheel is more about impl
15:36
<annevk2>
so Google does both Theora and H.264
15:36
<hsivonen>
annevk2: if the story is correct
15:37
<hsivonen>
annevk2: it doesn't make sense to me, though.
15:37
<jgraham>
annevk2: pointer?
15:37
<annevk2>
see inbox
15:37
<jgraham>
Oh, interesting
15:38
<jgraham>
It makes no sense to me either
15:38
<annevk2>
http://lists.whatwg.org/pipermail/whatwg-whatwg.org/2009-May/019992.html
15:38
<annevk2>
no sense how?
15:39
<Lachy>
annevk2, don't reinvent the wheel isn't about reusing implementations. It's about reusing existing solutions if they work
15:39
<hsivonen>
annevk2: if they don't think Theora is the kind of risk that Apple and Nokia claim, why not ship Theora only and switch YouTube to Theora?
15:40
<annevk2>
Lachy, I didn't say that
15:40
<Lachy>
then I don't understand what you meant by "Don't Reinvent the Wheel is more about impl"
15:41
<jgraham>
hsivonen: Possibly because H.264 would have better quality/bandwidth so they would save bandwidth costs exceeding the cost of licensing H.264
15:41
<Philip`>
hsivonen: Because it's lower quality and rarely used and doesn't have hardware support, perhaps?
15:41
<annevk2>
and then they'd need three versions of each video
15:41
<annevk2>
well, two, I guess
15:41
<jgraham>
hsivonen: Although it still requires that they ship a H.264 implementation with chrome
15:41
<hsivonen>
Philip`: rarely used would not apply if YouTube switched :-)
15:41
<jgraham>
annevk2: That seems like much less of a problem
15:42
<Philip`>
hsivonen: It would if you count the number of people who are familiar with the process of encoding to Theora :-)
15:42
<annevk2>
jgraham, hah, they get 20 hours of content every minute
15:42
<hsivonen>
annevk2: YouTube already has 3 per video: flv, H.264 and .3gp
15:42
<annevk2>
hsivonen, 3gp? isn't flv just streaming the H.264?
15:42
<jgraham>
I would be somewhat unsurprised if Youtube offered Theora to theora-only browsers
15:43
<hsivonen>
annevk2: there's traditional YouTube (flv), there's "HD" (h.264) and there's non-iPhone mobile (3gp)
15:43
<hsivonen>
3gp is MPEG-4 Simple Profile in small size
15:43
<Philip`>
The BBC iPlayer seems to have ten versions of some videos
15:43
<annevk2>
i thought they were in the process of converting the traditional stuff to h264
15:43
<annevk2>
but okay, I guess we'll see :)
15:44
<hsivonen>
annevk2: oh. I don't know about that.
15:44
<Philip`>
(http://beebhack.wikia.com/wiki/IPlayer_TV#Comparison_Table)
15:45
<Lachy>
hsivonen, there's actually standard quality, high quality, HD, plus the mobile phone versions
15:45
<Philip`>
(They probably don't get more than about a minute of video per minute, though)
15:45
<hsivonen>
it's amazing that there's enough compute power in the world to encode all those videos
15:46
<annevk2>
http://twitter.com/circa1977/statuses/1960304218 -- "... HTML will always be XML subset." look, someone is wrong on the internet!
15:46
<Lachy>
I assume YouTube also keeps the original uploaded videos around too, for any future conversions
15:46
<hsivonen>
I wonder if they are shipping both in order to demonstrate their ability to jettison h.264 to MPEG-LA
15:47
<hsivonen>
anyway, very cool that they'll ship theora
15:47
<Philip`>
hsivonen: It sounds like you seriously underestimate the amount of compute power in the world :-)
15:48
hsivonen
wonders what the carbon footprint of YouTube is
15:49
<Philip`>
hsivonen: If it's 20 hours of video per minute, you only need 1200 machines doing real-time conversion (and I guess you can do better than real-time on modern CPUs), multiplied by however many versions of videos you've got
15:49
<annevk2>
I guess H.264 won't be done via FFmpeg
15:49
<hsivonen>
Philip`: re-encoding the back catalog is a lot of minutes
15:51
<Philip`>
hsivonen: I wonder if they only bother converting the relatively popular videos to newer formats
15:51
<Philip`>
(I guess there are lots of videos that had 12 views a year ago and have never been looked at since)
15:53
<jgraham>
There are a rather large number of videos that are audio + a static picture (e.g. some copyright-infringing music "videos")
15:53
Philip`
would kind of guess that Google has lots of spare CPU capacity since most of its work will be limited by network I/O instead
15:53
<Philip`>
though I could be totally wrong
15:54
<Philip`>
jgraham: It's weird that a video sharing site is the most popular way to share music
15:54
<hsivonen>
I wonder if it's possible to recode lazily on demand
15:55
<Lachy>
Philip`, it's probably because there are no popular social networking sites built around the idea of publishing and listening to audio
15:57
<Philip`>
hsivonen: As in real-time encoding and streaming? That sounds like it'd be entirely incompatible with their usual architecture of serving videos as plain files from caching HTTP servers, so I guess it'd be quite a pain to do
16:00
<Lachy>
I doubt on-demand encoding would work given the number of simultaneous views they get across all their videos.
16:00
<Lachy>
but they would probably prioritise re-encoding of the back catalogue based on popularity or some other metrics
16:00
<Philip`>
Lachy: I assume the idea is they could recode the video on demand when somebody first views it, and then save the recoded output to send to any other viewers
16:01
<Philip`>
which would mean everybody would be able to see the recoded version of every video with no delay, without requiring a giant batch conversion before switching to the new format
16:05
<Philip`>
Streaming video seem to be the kind of thing that telecom companies care a lot about, so they want to invest in high-bandwidth high-reliability low-latency links and Quality of Service and multicast and all sorts of proprietary protocols and everything, but then consumers don't care about any of that and just download video files over HTTP and wait a few seconds until it's buffered enough to start playing smoothly
16:07
<hsivonen>
telcos care about QoS a lot more than justified by their customers' willingness to pay for it
16:15
Philip`
is not yet sure to what extent QoS is a real issue for the internet, and how much is just outdated Bellhead vs Nethead ideas
16:21
<mgrdcm>
hsivonen: one reason telcos care about it is to give their own VoIP traffic priority. they care about it within their own networks at least.
16:22
mgrdcm
used to work on big telco traffic quality monitoring software
16:23
<hsivonen>
mgrdcm: my point is that telcos fret about Erland formulas and circuit switching when the casual customer is happy with the Skype level of uncertainty
16:26
<mgrdcm>
hsivonen: for voice, though, i think customers' expectations of quality are higher when provided by a "phone company", even if it is over IP just like they'd be getting from vonage.
16:29
<hsivonen>
what does Vonage do for 911 calls?
16:30
<ezyang>
gsnedders: Houston, we have a problem.
16:30
Philip`
's university switched to VOIP last year, and every few months there are data network outages that cause the phone system to fail for several hours
16:31
ezyang
's university is switching to VOIP, but none of the students use the phones anyway...
16:32
<jgraham>
Philip`: Very locally that worked nicely for me since you couldn't hear anyone on the analouge phone that was in our office (it was too quiet)
16:32
<jgraham>
(and had a many-metre extension cable)
16:32
<Philip`>
(The phone system has actually been *less* reliable than the data network, since my building has a backup link for data but the phone is on a different VLAN that apparently won't work over the backup link)
16:33
<jgraham>
(so being able to hear people sometimes was overall a worthwhile gain)
16:34
<Philip`>
Analogue phones have the advantage that you don't have to wait a minute for them to reboot and acquire an IP address whenever there's a problem or whenever they get unplugged
16:35
<Philip`>
But when the VOIP phones work, they seem to work fine and they have new features and stuff, so I guess that'd good
16:35
<Philip`>
s/'d/'s/
16:37
<ezyang>
Do you have the feature where you can have voicemail emailed to you as an MP3?
16:39
gsnedders
notes that would mean paying the fees to encode an MP3
16:39
<Philip`>
ezyang: Not sure, but there is a web page where you can listen the voicemail as an .au file or something
16:40
<ezyang>
Basically the same
16:40
<Philip`>
(That's actually the only way I know how to listen to voicemail - I'm sure there must be some feature on the phone itself to do that, but I haven't figured it out)
16:40
<ezyang>
gsnedders: So, you know how we're optimizing tokenizer by globbing as many characters as we can? Well, it's actually kind of important to keep whitespace separated
16:40
<Philip`>
(But I've only had one voicemail message in the past year, so I haven't cared a lot)
16:40
<gsnedders>
ezyang: I know
16:41
<gsnedders>
ezyang: But only when we move in or out of the data state, IIRC
16:41
<Philip`>
I imagine you don't want to split on whitespace between words inside an element, because that'd hurt performance a lot
16:42
<ezyang>
Yeah... I'm trying to find which one is causing the test case to fail
16:42
Philip`
isn't sure how Python html5lib handles this (if it handles it at all)
16:42
<ezyang>
But if you have: <!DOCTYPE html><script> <!-- </script> --> </script> EOF
16:42
<ezyang>
(note space between </script> and EOF)
16:42
<gsnedders>
Philip`: It only gives WhitespaceCharacters when moving into the data state, IIRC
16:42
<ezyang>
The whitespace gets placed in <head>, but EOF gets placed in <body>
16:43
gsnedders
notes it is quite likely his memory is wrong
16:44
<ezyang>
Ok, this is definitely data' fault; my fix was just wrong
16:46
<ezyang>
What does HTML5 define as whitespace?
16:46
<gsnedders>
We should probably have a constant for that
16:46
<gsnedders>
but from memory: U+0009, U+000A, U+000C, U+000D, U+0020
16:47
<hsivonen>
gsnedders: yes
16:47
<ezyang>
Hmm. U+000B isn't one of them?
16:47
<gsnedders>
Nope.
16:47
<ezyang>
Ok, so, like, all of our code is wrong :-)
16:47
<hsivonen>
ezyang: not anymore
16:47
<gsnedders>
ezyang: Tok doesn't allow it
16:47
<ezyang>
Ah, that makes sense
16:48
<ezyang>
So, in TreeConstructer, we've got loads and loads of preg_match('/^[\t\n\x0b\x0c ]+$/'
16:48
<ezyang>
I should to a global search replace at some point
16:48
<ezyang>
In favor of a constant
16:48
<gsnedders>
Ah, in TreeConstructor
16:48
<ezyang>
Oh yeah, that's bugging me to
16:48
<ezyang>
We should rename the class
16:48
<ezyang>
*too
16:48
<gsnedders>
Also: remove pcre dependancy.
16:49
<ezyang>
I'll let you do that, since you're the resident expert :-)
16:49
<gsnedders>
Me!?
16:49
<hsivonen>
seems like you are using much higher-level constructs in the tokenizer than I am
16:49
<gsnedders>
hsivonen: This isn't tokenizer there.
16:49
<gsnedders>
hsivonen: We're about the same level in the tokenizer as Python
16:51
<ezyang>
gsnedders: InputStream takes care of \r, so I don't need to check for it, right?
16:51
<gsnedders>
ezyang: What about &#x0d; ?
16:52
<ezyang>
Oh, yeah, good point
16:53
<Philip`>
Doesn't &#x0d; get replaced in the tokeniser?
16:53
<ezyang>
Huh. Fixing that introduced three more fails
16:53
<gsnedders>
Philip`: dunno
16:54
<Philip`>
http://www.whatwg.org/specs/web-apps/current-work/multipage/syntax.html#tokenizing-character-references says replace 0x0D with U+000A
16:54
<gsnedders>
Oh well, OK
16:54
<gsnedders>
ezyang: You don't need to check for it.
17:00
<gsnedders>
ezyang: Where about are you working atm?
17:00
<ezyang>
I'm smashing tree builder bugs
17:00
<ezyang>
As long as TreeBuilderTest.php keeps working, you can do whatever you want to Tokenizer/InputStream
17:00
<ezyang>
Let me just commit and push my recent changes
17:00
<gsnedders>
Uh, yeah, I may end up touching TreeBuilderTest :)
17:01
<hsivonen>
http://en.wikipedia.org/w/index.php?title=HTML_5&diff=293104533&oldid=prev
17:01
<ezyang>
As long as it keeps working
17:01
<ezyang>
I'm not really editing that file per se. Just using it.
17:02
<ezyang>
Question: Why does the </cite> in <b>A<cite>B<div>C</cite>D *not* close the <div>?
17:03
<ezyang>
<cite> is not mentioned anywhere in the spec, so it should close the div...
17:03
<gsnedders>
"Please leave your sense of logic at the door, thanks!"
17:04
<ezyang>
No, this is w.r.t. the spec
17:04
<ezyang>
AFAICT, the spec mandates that <div> get closed
17:04
<ezyang>
But the test-case asserts differently
17:04
<ezyang>
Oh, my algo is wrong. Savvy
17:05
<Philip`>
"TreeConstructer.php"?!
17:05
<gsnedders>
Philip`: We didn't decide to call it that!
17:06
<ezyang>
Yeah, that's what the original was named
17:06
<ezyang>
and I didn't feel like renaming it without a bunch of tests first
17:06
<gsnedders>
Where has Jeroen gone? I haven't heard from him in a while
17:07
<Philip`>
You should rename it before anyone starts relying on it :-)
17:07
Philip`
is reminded of http://blogs.msdn.com/oldnewthing/archive/2008/05/19/8518565.aspx
17:07
<gsnedders>
Philip`: There are uglier things
17:07
<ezyang>
Philip`: aye-aye, sir
17:07
<Philip`>
gsnedders: Like the use of PHP?
17:07
<gsnedders>
Philip`: For example
17:08
<Philip`>
You really should have written it in Haskell instead
17:08
<ezyang>
Hahaha
17:08
<ezyang>
I mean, I'm totaly learning Haskell right now
17:08
<ezyang>
*totally
17:08
<gsnedders>
Philip`: Or Referer in HTTP
17:08
gsnedders
had learning Haskell on his to-do list for last summer
17:08
<gsnedders>
I never got to the end of the first item over the summer
17:09
<ezyang>
Haskell is totally worth skipping the rest of your list for.
17:09
<ezyang>
Catamorphisms yum!
17:10
<gsnedders>
So, InputStream tests…
17:11
<gsnedders>
What do I want to extend for the class?
17:11
<Philip`>
Python html5lib has a few you could steal, but I'm not sure how useful or appropriate or extensive they are
17:11
<ezyang>
Hmm... so it looks like foster parented elements don't get active formatting elements applied to them.
17:11
<ezyang>
Oh wait they do.
17:11
<gsnedders>
ezyang: I guess not HTML5_TestDataHarness…
17:12
<gsnedders>
ezyang: UnitTestCase?
17:12
<ezyang>
Hmm?
17:12
<ezyang>
Oh yeah, UnitTestCase is what you want
17:12
<ezyang>
that and $this->assertIdentical() are probably all you need
17:12
<ezyang>
maybe a setUp() or two
17:12
ezyang
does sit ups
17:12
gsnedders
hasn't used SimpleTest before
17:13
gsnedders
wishes it did code coverage…
17:14
<ezyang>
Does PHPUnit do code coverage these days?
17:14
<gsnedders>
Has for ages.
17:15
gsnedders
wishes there was a generic code-coverage tool for PHP that worked with any script
17:15
<ezyang>
YEah, it's called XDebug
17:15
<gsnedders>
That doesn't create anything nice to look at, just arrays
17:16
<ezyang>
Well, it's like a really simple PHP script to get some nice output
17:16
<ezyang>
(just someone has to write it)
17:16
gsnedders
doesn't remember it being overly simple
17:16
<gsnedders>
I started to write something, then realized it was more than I could be bothered to write
17:17
<ezyang>
Let me bang something out
17:19
<gsnedders>
Do we want one test per method?
17:20
<ezyang>
Yep.
17:20
<ezyang>
Pick expressive method names
17:26
<ezyang>
Done
17:26
<gsnedders>
Done what?
17:26
<ezyang>
With the highlight script
17:26
<gsnedders>
ah
17:26
<ezyang>
Let me stick in GitHub or something
17:26
<ezyang>
So you can try it out.
17:27
<Philip`>
Hmm, excellent, new versions of MySQL think width-140 with unsigned smallint width=80 should result in 18446744073709551556
17:27
<Philip`>
(compared to a much older version that thought it was some crazy like -60)
17:29
<ezyang>
Use case is to do the code coverage, and then save it as a .ser file
17:29
<ezyang>
http://github.com/ezyang/xdebug-highlight/tree/master
17:29
<ezyang>
I haven't tested it myself
17:31
<gsnedders>
ezyang: That doesn't break down into coverage % for functions, etc. :P
17:31
<hsivonen>
Hixie: how did http://dev.w3.org/html5/spec/Overview.html#declarative-3d-scenes end up in the spec?
17:32
<ezyang>
Oh, you want that info
17:32
<ezyang>
Yeah, then I officially don't care.
17:32
<ezyang>
Port our test-cases to PHPUnit or something
17:32
<gsnedders>
ezyang: Uh, that implies effort.
17:33
<ezyang>
You're working on html5lib. Come on, mate.
17:33
<gsnedders>
:P
17:33
<gsnedders>
ezyang: "Please leave your sense of logic at the door, thanks!"
17:33
<ezyang>
Anyway, I'm a SimpleTest dev, so I can fix problems and stuff.
17:33
<ezyang>
Also, SimpleTest has an experimental code coverage branch
17:34
<ezyang>
I've never used it though.
17:36
<Philip`>
hsivonen: There was some discussion about 3d <canvas> years and years ago, and someone mentioned X3D in there, so I guess that's what resulted in it being mentioned in the spec
17:37
<Philip`>
(Some of the X3D people seem to be misinterpreting it as stating that X3D is the official solution for embedding 3D in HTML5)
17:37
<hsivonen>
has any browser vendor shown interest in X3D?
17:38
<Philip`>
No, as far as I'm aware
17:38
<hsivonen>
Mozilla, Google and Opera have all tried something else
17:38
<Philip`>
All the current X3D support is just in plugins (or standalone applications), I believe
17:39
<Philip`>
but various X3D people want more integration with the browser
17:39
<Philip`>
e.g. linking it directly to the DOM inside an XHTML page, like with SVG
17:43
<Philip`>
(Maybe it could be useful in the same way that having both SVG and 2d <canvas> is useful, since they're appropriate for different situations)
17:44
<Philip`>
(Alternatively, maybe it could be a waste of effort in the same that having both SVG and 2d <canvas> is a waste of effort, since there's significant overlap)
17:48
<ezyang>
Subtle: if we reconstruct an active formatting element, *that* becomes the foster parent. I wonder where the python implementation special cases this.
17:52
<Dashiva>
So if you build a paved road, and later a few lonely cows start walking on it... does it even make sense to call it a cowpath?
17:55
<Philip`>
What if you can't see any cows at all, but the road is covered in cow pats?
17:56
<ezyang>
Is it a cow if no one sees it?
17:56
<Philip`>
ezyang: Yes
17:58
<Dashiva>
Well, I'm just wondering because there seem to be a few people saying "<existing spec feature> is a cowpath because some people use it"
18:04
<gsnedders>
Dashiva: But cows are cute!
18:04
<gsnedders>
We must keep them!
18:08
<Philip`>
gsnedders: That's the main reason I use Gentoo
18:08
<Philip`>
"emerge moo" prints a nice picture of the mascot, Larry the Cow
18:10
<Dashiva>
Philip`: You linked that bytes-to-encoding image again, which reminded me that you were going to make one with a more sensible y axis :)
18:10
<Philip`>
Was I?
18:11
<Dashiva>
Logarithmic in the percentage of non-working sites, I think
18:11
<Philip`>
You mean like http://philip.html5.org/data/encoding-detection-loglog.svg or something?
18:14
<Dashiva>
Yeah, that
18:15
<Dashiva>
I guess I just forgot you had made it :)
18:16
<Philip`>
How could you forget? It wasn't even 15 months ago
18:16
<Philip`>
(I suppose I do have the distinct advantage of being able to see the directory listing...)
18:17
<Dashiva>
Maybe I assumed you would've linked that one if it existed, even though it's just my personal taste saying it's better
18:17
<Philip`>
The other graph does a better job of misrepresenting the data so that 512 bytes looks like a sensible figure to choose
18:18
<Dashiva>
One thing I've wondered is, does it really make a difference to wait until 1024 bytes?
18:18
<Dashiva>
Isn't that usually within the first packet anyhow?
18:18
<Philip`>
Usually, but not close to always
18:19
<Philip`>
http://krijnhoetmer.nl/irc-logs/whatwg/20090528#l-1107
18:20
<Philip`>
(Packets are usually a bit under 1500 bytes, I think)
18:21
<Dashiva>
Hmm, so if the header is 400+ you might run past the first packet
18:21
<gsnedders>
ezyang: Right, so I think we basically need to have our own UTF-8 decoder, but we can skip it if we have an up-to-date copy of PCRE with Unicode or a sane iconv, _and_ we have PCRE.
18:22
<ezyang>
Sounds like it. hsivonen has one
18:22
<gsnedders>
I have one too.
18:22
<gsnedders>
I have an MIT licensed one, which is more useful for html5lib than hsivonen's
18:23
<Dashiva>
My internet connection is too fast for me to be able to care about the difference in one and two packets, I suppose :)
18:25
<gsnedders>
ezyang: The only reason we need PCRE always for when we don't need to use our own UTF-8 decoder is what is currently line 130 of inputstream
18:25
<ezyang>
I've seen. It's quite an impressive regex you have there.
18:26
<gsnedders>
Yeah. It was fun to write. </sarcasm>
18:26
<Philip`>
Dashiva: Latency matters more than bandwidth, and I assume your internet connection still suffers from latency when accessing data on the other side of the world
18:26
<Philip`>
although I don't how much it actually matters in reality
18:26
<Philip`>
because I don't really have any idea how TCP works
18:26
<Philip`>
(Will it send lots of response packets at once, rather than doing the first and waiting for an ACK?)
18:26
<gsnedders>
ezyang: I used to require PCRE with UTF-8 support there, but that says any string containing low surrogates is invalid, which doesn't work when we want to detect them.
18:27
<gsnedders>
(It's also cheaper not having to decode it)
18:27
<ezyang>
Awesome.
18:27
<Dashiva>
Philip`: It will send several, yes
18:28
<Philip`>
I guess this has something to do with the default window size
18:28
<sr0unet>
Hello.
18:28
<gsnedders>
ezyang: (Most of the test case failures I have so far come from using iconv!)
18:28
<Dashiva>
Yeah
18:28
<sr0unet>
Is this a WPF channel ?
18:28
<gsnedders>
s/most/all/
18:28
<Dashiva>
That's how far ahead it's willing to go, as I recall
18:28
<Dashiva>
sr0unet: What's WPF? Wikipedia foundation?
18:28
<sr0unet>
^^
18:28
<Dashiva>
Then no
18:28
<ezyang>
As in, we need to emit parse errors and iconv doesn't do that?
18:28
<Philip`>
Walrus protection fund?
18:28
<sr0unet>
WPF C#
18:28
<gsnedders>
ezyang: No
18:29
<gsnedders>
ezyang: Replace invalid sequences with U+FFFD
18:29
<Dashiva>
This is where HTML happens
18:29
<Philip`>
I don't think anyone using an HTML5 parser library is actually going to care about parse errors, so it'd be much easier to not implement them :-)
18:29
<gsnedders>
Philip`: But we already have :P
18:30
Philip`
should add a no-error mode to Python html5lib so it doesn't have to go through all the bother of tracking line numbers and suchlike
18:33
<ezyang>
aha
18:33
<ezyang>
Philip`: Actually, people surprisingly do
18:43
<gsnedders>
Wait, what…
18:43
<gsnedders>
DOM Level 2 Events uses "thru"
18:45
<Dashiva>
To the Errata-mobile
18:48
<gsnedders>
How do you get each member of a NodeList in ECMAScript?
18:48
<gsnedders>
My out of practiceness of JS shows
18:51
<gsnedders>
Can you not set an event handler on DOMActivate?
18:52
<ezyang>
Hey, question for y'all: when the spec says 'Process the token using the rules for the "in head" insertion mode.', do they implicitly want me to temporarily add <head> to the top of the stack?
18:52
<Dashiva>
gsnedders: .length and iteration?
18:52
<ezyang>
This is section 9.2.5.10
18:52
<gsnedders>
http://hixie.ch/tests/adhoc/dom/events/DOMActivate/001.html makes it seem you can, but Op 10 fails
18:52
<Dashiva>
DOMActivate is part of mutation, isn't it?
18:54
<gsnedders>
Oh, wait. I realize what I'm doing wrong.
18:56
<gsnedders>
Is it bad I more or less permanently have HTML 5 open?
18:56
<ezyang>
Hmm, it looks like the spec doesn't actually want that
18:56
<Dashiva>
Philip`: What's your comment to being labeled as not part of the whatwg crowd?
18:57
<hsivonen>
http://lists.w3.org/Archives/Member/tag/2009May/0084.html
18:59
Dashiva
misses being member of a Member so he could read Member mails
18:59
<Philip`>
Dashiva: Depends on whether that comment came from a cool person i.e. a WHATWG member, or someone else
19:01
<Dashiva>
It was our friend Shelley
19:01
<gsnedders>
Does 'a, img[usemap], video[controls], audio[controls], label, input:not([type=hidden]), button, select, textarea, keygen, details, datagrid, bb, menu[type=toolbar]' match all interactive elements?
19:04
<ezyang>
Ok, I think I see a bug in the test-cases
19:04
<Philip`>
Dashiva: Hmm, where was that?
19:04
<ezyang>
<head></html><meta><p> should result in <meta> being placed in <body>, since </html> pops us to AFTER_AFTER_BODY
19:05
<ezyang>
Oh shoot, I'm in the wrong mode
19:05
<Dashiva>
http://lists.w3.org/Archives/Public/public-html/2009May/0537.html
19:05
<ezyang>
Oh no, it's the same behavior
19:05
<ezyang>
Awesome. I think this is legit. Anyone else want to verify?
19:06
<hsivonen>
http://lists.w3.org/Archives/Public/www-tag/2009May/0123.html has interesting bits
19:07
<ezyang>
The question is does <head></html> result in AFTER_AFTER_BODY or IN_HEAD mode
19:07
<Philip`>
Dashiva: Oh, right
19:08
<hsivonen>
ezyang: V.nu may have a bug in that case
19:08
<Dashiva>
ezyang: What's the non-legit behavior seen elsewhere?
19:09
<ezyang>
The python, V.nu, and test-case think that tihs should result in IN_HEAD
19:09
<ezyang>
I think it should be AFTER_AFTER_BODY
19:09
<Dashiva>
Okay
19:12
<Dashiva>
So the </html> is "in head", becomes "after head", becomes "in body", becomes "after body", becomes "after after body". Then the <meta> is sent to "in body", and then sent to "in head"
19:13
<ezyang>
Dashiva: yep.
19:14
<ezyang>
But the "in body" case doesn't push the head_pointer onto the stack, so meta ends up in body, not html
19:14
<Dashiva>
But the <meta> ends up in the head
19:14
<ezyang>
s/html/head/
19:14
<ezyang>
If meta should end up in head, I think we have a spec bug.
19:14
<Dashiva>
Oh, you mean in the code?
19:14
<ezyang>
The Python/V.nu code puts it in the head. I think spec puts it in body.
19:14
<Dashiva>
Not that I can see
19:15
<ezyang>
It's subtle. When you're "in body", it passes it on to "in head"
19:15
<Dashiva>
Here's how I see it in the spec: You start in "after after body". You follow the step "anything else" which sends you to "in body"
19:15
<Dashiva>
There you follow A start tag token whose tag name is one of: "base", "command", "link", "meta" ... which takes you to "in head"
19:15
<ezyang>
But the top of the stack is still <body>, not <head>
19:15
<ezyang>
yep and yep
19:15
<Dashiva>
Ah, I see what you mean
19:15
<ezyang>
We are in agreement.
19:16
<hsivonen>
ezyang: at some point, the spec said that </body> and </html> were ignored in head
19:16
<ezyang>
Yep.
19:17
<ezyang>
So, assuming that the spec change is correct, this test-case needs to change
19:17
<hsivonen>
yes
19:18
<ezyang>
Ok, will do.
19:18
<ezyang>
I'll also patch the Python implementation for free :-)
19:18
<Dashiva>
But the intent is still that <meta> should always end up in head, isn't it?
19:19
<Dashiva>
That is, it's a spec bug too
19:19
<ezyang>
Well, we have test-cases that very specifically say <meta> should end up in <body>
19:19
<Dashiva>
Oh
19:19
<Dashiva>
Microdata, of course
19:19
<Dashiva>
I had things mixed up
19:21
<hsivonen>
Dashiva: it predates microdata. IE and Opera compat
19:21
<hsivonen>
the parsing spec was *not* changed for microdata
19:21
hsivonen
waves to log readers
19:21
<Dashiva>
But remembering microdata told me which way is the intended one :)
19:22
<ezyang>
Browsers are... special.
19:22
<ezyang>
Yup yup
19:22
<Dashiva>
hsivonen: It's a bit worrisome if anyone were to take anything I say as reasonable
19:22
jgraham
wonders if Philip` has a clever way of making parseerrorless html5lib significantly faster whilst still allowing the possibility of reporting parse errors
19:23
gsnedders
notes the perf. hit was fairly low even in PHP
19:23
<Dashiva>
jgraham: Would it be parseerrorless if it had that capability?
19:23
<Philip`>
jgraham: Depends on whether 5% counts as significant
19:23
<hsivonen>
jgalvez_: does python compile away empty methods that could be filled in by a subclass but aren't at the time of execution?
19:24
<gsnedders>
Does 'a, img[usemap], video[controls], audio[controls], label, input:not([type=hidden]), button, select, textarea, keygen, details, datagrid, bb, menu[type=toolbar]' match all interactive elements?
19:24
<ezyang>
"two-heads-are-not-better-than-one" HAHAHA
19:24
<hsivonen>
s/jgalvez_/jgraham/
19:24
<jgraham>
hsivonen: Assuming you meant me, no idea but I doubt it
19:25
<jgraham>
Philip`: How?
19:26
<hsivonen>
jgraham: I'm refactoring code assuming that HotSpot does that
19:26
<Dashiva>
hsivonen: So tag is getting involved in the RDFa side of things?
19:26
<jgraham>
ezyang: Ah, the delicate touch of mpilgrim
19:26
<Philip`>
jgraham: By not bothering to keep track of line numbers
19:26
<hsivonen>
Dashiva: looks like it might
19:26
<Philip`>
jgraham: (I think it saved roughly that much when I last checked)
19:27
<jgraham>
Philip`: So just an if (parse_errors): calculate_line_numbers type thing
19:27
<Philip`>
jgraham: Something like that
19:27
<jgraham>
It would be nice to cut all the method calls to parseError()
19:27
<jgraham>
But that is harder I guess
19:28
<Dashiva>
At least the talk seemed rather calm and reasonable
19:28
<ezyang>
jgraham: Ah, so this is mpilgrim's baby :-)
19:28
<Dashiva>
No flaming rhetoric from the get-go :)
19:28
<jgraham>
ezyang: mpilgrim tried making a html5lib based validator
19:28
ezyang
.oO{ghetto}
19:28
<jgraham>
Which never really went anywhere
19:28
<ezyang>
Ah.
19:29
<ezyang>
So, sorta like V.nu, except vapourware?
19:29
<jgraham>
ezyang: sort of
19:30
<Philip`>
ezyang: Not really, since there wasn't any vapour
19:30
<Philip`>
Just some code that checked a few conformance requirements
19:30
<ezyang>
Aha.
19:31
<ezyang>
(in other news, test-case fix pushed. Rubyistas and Javaistas, fix yer code!)
19:31
<ezyang>
Maybe I should fix the ruby impl too.
19:33
<gsnedders>
When was DOM 3 XPath support added to browsers?
19:42
<hsivonen>
gsnedders: March 2002
19:43
<hsivonen>
gsnedders: http://bonsai.mozilla.org/cvslog.cgi?file=mozilla/dom/public/idl/xpath/nsIDOMXPathEvaluator.idl&rev=HEAD&mark=1.3
19:43
<hsivonen>
dunno about other code bases
19:43
<ezyang>
Ugh, fixing the ruby impl increased the fail count by two
19:52
<hsivonen>
http://twitter.com/jdowdell/status/1961148437
19:53
<ezyang>
Is there any particular reason there's a smattering of test-cases that have error messages like "Unexpected end tag (). Ignored."?
19:54
<ezyang>
(four, specifically)
19:58
<jwalden>
judging by this channel, jd is an extremely effective troll
20:01
<Hixie>
don't confuse people being entertained with people being enraged :-)
20:02
<jwalden>
lots of talk for not that much entertainment, to me :-)
20:02
<ezyang>
I'm not poking the ruby implementation any more (although I managed to introduce 10 more failing cases.) >:-(
20:24
<Wolfman2000>
Morning/afternoon. www.pumpproedits.com <-- I present, the newest member of the HTML5 family.
20:29
<ezyang>
tests1.dat -> full passes with PHP implementation
20:29
<ezyang>
moving on to tests2.dat
20:30
<Wolfman2000>
...testing? What type of testing?
20:31
<ezyang>
Woflman2000: the PHP port of html5lib
20:33
<Wolfman2000>
Will a python library be required?
20:34
<ezyang>
Nope. That's the port of a port.
20:34
<Wolfman2000>
...I'll get clarification on that later
20:34
<ezyang>
erm, *point of a port
20:35
<ezyang>
(port of a point? port of a sherry? The mystery deepens)
20:35
<gsnedders>
What about the side of a port, then?
20:35
<Wolfman2000>
I'll rephrase the question
20:36
<Wolfman2000>
will a Python port be required?
20:36
<Wolfman2000>
to be built
20:36
<ezyang>
We already have a Python port
20:37
<Wolfman2000>
*nods*
20:37
<ezyang>
(I suppose, if the Python impl was first, it's not technically a port...)
20:37
beowulf
has an image of snakes weighing anchor
20:38
<ezyang>
We can call it starboard, or something
20:40
<Philip`>
We can just call it html5lib
20:40
<Philip`>
(All the others are pretenders)
20:41
<ezyang>
Ok, another test-case bug (I think)
20:42
<ezyang>
nvm.
20:43
<mpilgrim>
http://www.youtube.com/html5 works in chrome 3.0.182.3
20:43
<Wolfman2000>
...and my copy of Firefox won't work on that
20:44
<Wolfman2000>
3.0.10 for the record
20:44
<mpilgrim>
http://commons.wikimedia.org/wiki/Category:Ogg_video works in chrome 3.0.182.3 too
20:48
<mpilgrim>
mark@atlantis:~% mp4creator -list google_main.mp4
20:48
<mpilgrim>
Track Type Info
20:48
<mpilgrim>
1 audio MPEG-4 AAC LC, 58.049 secs, 125 kbps, 44100 Hz
20:48
<mpilgrim>
2 video H264 Baseline⊙2, 57.599 secs, 499 kbps, 480x270 @ 29.983159 fps
20:48
<mpilgrim>
(google_main.mp4 is the video on http://www.youtube.com/html5 )
20:48
<mpilgrim>
wikimedia commons has theora video with vorbis audio in an OGG container
20:50
<Wolfman2000>
mpilgrim: the youtube.com/html5 thing is NOT working on Firefox 3.5 beta 4 on Mac OS X
20:50
<Wolfman2000>
the video is not playing
20:51
<scherkus>
the youtube site uses h264/aac, which as of right now is only supported in chrome 3.0.182.3 and safari 4 beta
20:51
<mpilgrim>
i am aware of that
20:51
<Wolfman2000>
scherkus: thank you
20:51
<mpilgrim>
i'm testing google chrome right now
20:52
<mpilgrim>
http://wearehugh.com/public/2006/12/20061225-large.mp4 plays in google chrome 3.0.182.3
20:52
<mpilgrim>
stats on that:
20:52
<mpilgrim>
Track Type Info
20:52
<mpilgrim>
1 video H264 Main⊙5, 142.842 secs, 751 kbps, 640x480 @ 29.970177 fps
20:52
<mpilgrim>
2 audio MPEG-4 AAC LC, 142.762 secs, 0 kbps, 48000 Hz
20:52
<mpilgrim>
so that's H.264 Main Profile video, AAC-LC audio, in an MP4 container
20:53
<mpilgrim>
i don't have any H.264 High Profile video handy
20:53
<Wolfman2000>
...wonder when Firefox will support h264/aac then
20:54
<hsivonen>
mpilgrim: is the h.264 part open source?
20:54
<Rik|work>
btw, openvideo on dailymotion now works in safari 4
20:54
<mpilgrim>
i'm assuming it's ffmpeg
20:54
<mpilgrim>
as discussed a few days ago
20:55
<hsivonen>
Rik|work: with XiphQT
20:55
<hsivonen>
mpilgrim: interesting considering licensing
20:55
mpilgrim
wonders if google chrome for mac could bundle XiphQT
20:55
<hsivonen>
Rik|work: with XiphQT?
20:55
<Rik|work>
hsivonen: no, with mp4
20:55
<mpilgrim>
bbiab
20:55
<hsivonen>
Rik|work: also interesting
20:56
<Rik|work>
hsivonen: but they're not using <source>, they do browser sniffing
20:56
<ezyang>
Ok, this algorithm is confusing me
20:56
<hsivonen>
:-(
20:56
<Rik|work>
I'm talking with the devs (I'm french) and they answered they have many reasons for not using that
20:56
<ezyang>
Otherwise, set node to the previous entry in the stack of open elements and return to step 2." and then "Initialize node to be the current node (the bottommost node of the stack). "
20:56
<Rik|work>
one of them was lack of support in a previous firefox beta
20:57
<ezyang>
Does the second step mean that we pop nodes node is the current node, or what?
20:57
<ezyang>
Or does the spec actually mean I should go to step 3?
20:58
<ezyang>
Hmm, looks like I somehow fixed it. Nevermind
21:01
<Groovy>
you noticed google.com stopped being <!doctype html> ?
21:01
<Wolfman2000>
I never noticed them using it to begin with.
21:01
<ezyang>
Huh. We don't have PLAINTEXT support in our tokenizer. wtf
21:03
<ezyang>
Huh. The spec doesn't handle PLAINTEXT
21:03
<Groovy>
Wolfman2000: yea, about from beginning of may to now they went <!doctype html>
21:03
<ezyang>
Uhh... Hixie? Any comments?
21:03
<Wolfman2000>
...they must be secretly in bed with Microsoft
21:04
<Wolfman2000>
HTML5 doesn't work for IE unless you use javascript to add the elements to the DOM
21:07
<gsnedders>
ezyang: It's an optimization.
21:07
<ezyang>
Huh?
21:07
<ezyang>
Sorry, I don't follow.
21:08
<gsnedders>
ezyang: PLAINTEXT state means everything stays in the data state
21:08
<gsnedders>
ezyang: We handle it by special casing other content models
21:08
<ezyang>
Uhm, no we don't
21:08
<gsnedders>
ezyang: Yes we do.
21:08
<gsnedders>
ezyang: We do nothing for it. In it, you will never move out of the data state.
21:08
<ezyang>
If we did, this test would pass: <table><plaintext><td>
21:09
<ezyang>
the tokenizer is happly continuing parsing even after being put in PLAINTEXT content model
21:09
<ezyang>
*happily
21:09
<gsnedders>
We should be in the data state after it
21:10
<ezyang>
Right. And we never check for mode PLAINTEXT
21:10
<gsnedders>
We don't need to.
21:10
<ezyang>
REally?
21:10
<gsnedders>
Hence we always stay in the data state.
21:10
<gsnedders>
ezyang: Look at the spec.
21:10
<gsnedders>
ezyang: We only move out of it after checking state is PCDATA, RCDATA or CDATA
21:11
<gsnedders>
The tokenizer looks fine to mee.
21:11
<gsnedders>
*me
21:12
<ezyang>
Oh, well, that's because PLAINTEXT content model is not getting set.
21:12
<gsnedders>
:D
21:12
<gsnedders>
See, the code I've been working on works. :P
21:13
<ezyang>
Yep.
21:13
<ezyang>
You're right
21:13
<gsnedders>
Of course.
21:15
<ezyang>
This return the content model business is obnoxious
21:15
<ezyang>
I think I'm going to refactor it out.
21:42
<ezyang>
Ok, so if I have <p><b></p> <p> does the space get wrapped by <b>?
21:43
<ezyang>
I'm leaning towards yes, althought the test case claims it doesn't.
21:43
<ezyang>
Also, the Python implementation agrees with me
21:43
<ezyang>
So does the Ruby implementation
21:44
<ezyang>
(could it be? could ezyang have actually found a real test-case bug?
21:46
<ezyang>
It could be a spec bug, since what then happens is we stick the <p> inside the <b><i><u> tags, which is not what active formatting elements is supposed to do.
21:47
<ezyang>
Any comments?
21:48
<jgraham>
ezyang: WDVND?
21:48
<ezyang>
come again?
21:49
<jgraham>
(What does Validator.nu do?)
21:49
<ezyang>
Oh, yeah
21:50
<ezyang>
V.nu does the test-case behavior
21:51
<ezyang>
I guess I could figure out what they did different.
21:53
<jgraham>
ezyang: Henri is generally more up to date than html5lib. Although there is a case with whitespace that he preemptively changed anticipating a spec change that hasn't happened
21:54
<ezyang>
Yep, that seems like it.
21:54
<ezyang>
hsivonen is not reconstructing active formatting elements when there's whitespace
21:55
<ezyang>
The spec hasn't changed w.r.t. yet, though
21:56
<ezyang>
I thought tokenizing/treebuilding was pretty stable now?
22:00
<ezyang>
"Update the tree building tests with frameset-ok, new AAA and (not yet in spec) WebKit-style foster-parenting"
22:00
<hsivonen>
if something depends on whitespace in 'in body', it's likely a bug
22:00
<ezyang>
So it's the webkit thing again :-)
22:00
<hsivonen>
oh, you meant foster parenting
22:00
<hsivonen>
that's deliberate
22:01
<ezyang>
So, I understand that the spec may not be at the point that V.nu is
22:01
<ezyang>
But having tests that the spec deliberately fails really doesn't help me out, if I'm trying for 0 fails
22:01
<hsivonen>
ezyang: when I checked that in, Hixie's IRC statements hinted at spec going that way
22:02
<ezyang>
Anyway, it's whitespace "in body"
22:03
<ezyang>
Specifically, the difference is V.nu checks if a character token is whitespace, and if it is, doesn't reconstruct active formatting elements
22:03
<hsivonen>
ezyang: not while foster parenting?
22:03
<ezyang>
Nope.
22:03
<ezyang>
No foster parenting at all.
22:03
<ezyang>
You changed the test case, however, in commit 535a1040e2df
22:03
<hsivonen>
weird
22:04
<ezyang>
for which I just posted the summary
22:04
<hsivonen>
wait, have the test cases moved to hg?
22:04
<hsivonen>
I've been checking in to svn still
22:04
<ezyang>
Yes.
22:05
<hsivonen>
ouch
22:05
<ezyang>
Does Google Code magically transfer it?
22:06
<ezyang>
I mean, there's no way to tell if there was an svn repository from Google Code website anymore
22:06
<jgraham>
ezyang: Dunno. I had to do the initial migration manually though
22:06
<ezyang>
Oh hey, that's the last commit you made, according to mercurial
22:07
<ezyang>
(re hsivonen)
22:07
<ezyang>
jgraham: aha
22:07
<ezyang>
When did the migration occur?
22:07
<hsivonen>
ezyang: you'll save a lot of EOF grief if you get my changes from this week from svn
22:08
<ezyang>
But I don't even remember where the svn repository *is*
22:08
<hsivonen>
I don't either. I'm not on the computer that has my html5lib sandbox
22:09
<ezyang>
Anyway, not turning off access to the subversion repos seems... poor.
22:09
<ezyang>
I'm not sure how we can easily merge your changes back in.
22:09
<jgraham>
ezyang: https://html5lib.googlecode.com/svn/trunk/
22:09
<ezyang>
Awesome
22:09
<jgraham>
It seems like svn should become read-only
22:11
<jgraham>
(I mean ideally, not that it does)
22:13
<ezyang>
AWesome, we only lost one commit
22:13
<ezyang>
(maybe two, if one of those was in a branch)
22:14
<ezyang>
hsivonen: you can probably take its diff and apply to the new hg repository
22:14
<hsivonen>
I haven't committed to branches
22:14
<hsivonen>
ezyang: I'll do that when I get to my work computer
22:15
<ezyang>
Odd. 1303 has nothing in it
22:15
<ezyang>
Great.
22:16
<ezyang>
Holy crap that's a big change
22:17
<ezyang>
Oh, these are all tokenizer updates
22:17
<ezyang>
Hmm... we already pass all of the tests... uh oh
22:17
<ezyang>
Anyway, hsivonen, you maintain that Hixie is going to change the spec to follow V.nu behavior?
22:20
<hsivonen>
ezyang: I intend to keep pushing Hixie that way :-)
22:20
<Wolfman2000>
...huh. The <header> tag has changed functionality?
22:20
Wolfman2000
catches up on the blog reading
22:20
<Wolfman2000>
...am I really only allowed to use <header> with <h#>?
22:20
<hsivonen>
ezyang: the tokenizer changed to discard tag token if EOF happens inside tag
22:20
<ezyang>
Ok. Do you mind if I separate out that test to its own file?
22:20
<ezyang>
Aha.
22:21
<hsivonen>
ezyang: that would be good
22:25
<jgraham>
Wolfman2000: You can only use <hgroup> with <hx> <header> you can use with anything
22:25
<ezyang>
Done.
22:25
<ezyang>
I should document this somewhere
22:26
<Wolfman2000>
ok
22:31
<ezyang>
hsivonen: I assume test 44 is another one of those?
22:31
<ezyang>
(in tests2.dat)
22:32
<ezyang>
This one is <p><b><i><u></p>\n<p>X
22:33
<hsivonen>
ezyang: if that one deviates from spec, it's a bug
22:35
<ezyang>
It's the same thing (note that in my post \n was meant to be interpreted as an scape)
22:35
jgraham
wonders how to stop hg merge from merging the python3 directory into the python directory
22:35
<ezyang>
V.nu doesn't reconstruct active formatting elements on \n; everyone else does
22:35
<hsivonen>
ezyang: sounds like a bug in V.nu
22:35
<ezyang>
You could ask #mercurial about it
22:35
<ezyang>
But... \n is whitespace
22:36
<hsivonen>
ezyang: I need to study what code I have written
22:36
<ezyang>
ok
22:37
ezyang
thinks "Wow, it's not often I get three reference implementations :-)"
22:39
<ezyang>
THIRTY-SIX!
22:39
<ezyang>
(failures, that is)
22:41
<ezyang>
Ok, now I need to implement the fragment algorithm
22:50
<ezyang>
I think that's enough html5lib hacking for today
22:50
<ezyang>
Test suites 1-3 are now passing