06:25
<Hixie>
i'm no longer checking html5 for validity because validator.nu won't let me check it -- says it's too long :-(
06:25
<Hixie>
hsivonen: any chance you could increase the limit again :-)
07:21
<othermaciej>
annevk42: are you around?
08:08
<hsivonen>
Hixie: I'll fix when I get to my work computer
08:09
<Hixie>
sweet :-)
08:09
<hsivonen>
I hadn't realized RDFa had this random restriction: http://www.w3.org/TR/rdfa-syntax/#col_Metainformation
08:10
<hsivonen>
Did I read right that Shane called UAs that lowercase names "legacy"?
08:11
<othermaciej>
hsivonen: well they haven't updated to XHTML2 yet
08:11
<othermaciej>
so obviously
08:12
<Philip`>
hsivonen: Which restriction are you referring to?
08:12
<hsivonen>
Philip`: not generating triples for rels not on that list
08:13
<Philip`>
hsivonen: Ah
08:13
<Philip`>
Seems a really bad idea in terms of forward-compatibility
08:14
<Philip`>
(At least one implementation (rdfQuery) throws a fatal error if you have a rel keyword that's not on the list, which can't be good)
08:14
<hsivonen>
Philip`: are you implying compliance isn't good? :-)
08:16
<Philip`>
hsivonen: Compliance doesn't require implementations to throw fatal errors, just to ignore the value, as far as I can tell
08:16
<hsivonen>
ah
08:33
mhausenblas
wondering what in http://www.w3.org/TR/rdfa-syntax/#col_Metainformation is random to hsivonen
08:36
<hsivonen>
mhausenblas: the admittance of non-HTML4 relations seems arbitrary
08:36
<mhausenblas>
ahm. and if you compare that to http://wiki.whatwg.org/wiki/RelExtensions - this is good practice, then, right? ;)
08:37
<hsivonen>
mhausenblas: license is there. also role. nofollow isn't there
08:37
<mhausenblas>
if nofollow isn't there then it's a bug, IMHO
08:37
<mhausenblas>
I'll check and bring it up at next telecon, thanks
08:38
<hsivonen>
mhausenblas: the whatwg wiki doesn't make future-limiting processing restrictions
08:38
mhausenblas
is not in the position to judge, but I personally like and prefer the 'open' vocabulary world
08:39
Philip`
will refrain from making any suggestions on handling rel="alternate stylesheet" similarly to how HTML5 microdata handles it (as a single token)
08:39
<hsivonen>
a wiki page seems more 'open' than a list created by a TF
08:39
<mhausenblas>
experience tells that even if you allow people to come up with their own voc, over time the economics enforce few vocs per topic to survive
08:40
Philip`
also refrains from commenting on rel="STYLESHEET" etc, which seemingly won't be accepted by an RDFa processor
08:40
<mhausenblas>
hsivonen: open in the sense of everyone can publish everywhere everything (why a single Wiki?)
08:41
<hsivonen>
mhausenblas: to avoid the step where you have overlapping vocabularies
08:41
<mhausenblas>
hsivonen: in my world this does not happen as each voc lives in its own namespace
08:41
<Philip`>
Regardless of where valid rel values are defined, they're going to be defined somewhere outside of RDFa, and so if RDFa defines it own parallel list then it's going to get out of sync at some point
08:41
<Philip`>
s/it/its
08:42
<mhausenblas>
I guess the Java and other communities have experienced this for years and it seems to work fine
08:42
<mhausenblas>
I have absolutely no idea why people oppose namespaces (note: I'm not talking about XMLNS per se)
08:45
mhausenblas
BRB, gotta grab some coffee ;)
08:54
<hsivonen>
mhausenblas: I meant semantically overlapping
08:55
<hsivonen>
Philip`: already got out of sync
08:55
<hsivonen>
(even for XHTML2!)
08:55
<mhausenblas>
hsivonen: ok. but I don't see the problem there
08:56
<mhausenblas>
take for example Goolge. they decided not to reuse the already available vocs such as FOAF, etc.
08:56
<mhausenblas>
fine. no problem. see for example http://iandavis.com/blog/2009/05/googles-rdfa-a-damp-squib#comment-1407
08:57
<hsivonen>
mhausenblas: some people who seem to care more about RDF than about RDFa cheerleading see it as a problem
08:57
<hsivonen>
mhausenblas: http://www.jenitennison.com/blog/node/104 (incl. comments)
08:58
<mhausenblas>
hsivonen: not sure if I understand this. RDFa is just an RDF serialisation
08:59
<mhausenblas>
yes, I know Jeni's post and have left my 0.02€ at http://lists.w3.org/Archives/Public/public-lod/2009May/0095.html
08:59
<Philip`>
It sounds like Google is using it as a serialisation of a non-RDF data format, rather than of RDF
08:59
<Philip`>
(and they're using microformats as a different serialisation of the same internal non-RDF data format)
08:59
<mhausenblas>
Philip`: there you might be right ;)
09:00
<hsivonen>
Philip`: indeed, but it's stille cheered as a success for RDFa. go figure
09:00
<mhausenblas>
the thing is: I'm interested in the Web of Data where data comes from whatever nicely machine-processable format
09:01
<hsivonen>
Philip`: anything that has the slightest appearance of using RDFa is cheered as a success, regardless of conformance or RDF model alignment
09:01
<Philip`>
hsivonen: Kind of like cheering when Google starts using "<!doctype html>" on some of their pages? :-)
09:02
<hsivonen>
Philip`: cheering for just the privacy page is embarrassing
09:02
<othermaciej>
it does seem like there is some degree of fanboyism around RDFa
09:02
<hsivonen>
Philip`: but at least it validated as HTML5 instead of just using the doctype
09:02
<othermaciej>
like excitement about its superficial aspects instead of about the goals it is meant to achieve
09:03
<mhausenblas>
hsivonen: IIRC the microformats community even had a party celebrating it, but this is not the point (and I don't care)
09:03
<othermaciej>
I'll be excited when Web pages use new HTML5 tags/attributes/APIs
09:04
<hsivonen>
mhausenblas: celebrating what?
09:04
<othermaciej>
the doctype is not remotely exciting
09:04
<mhausenblas>
hsivonen: that Google supports microformats
09:04
<Philip`>
It seems there's some degree of fanboyism around every technology
09:04
<hsivonen>
mhausenblas: wow
09:05
<mhausenblas>
to be fair Gentlemen, I think each and everyone of us who has invested couple of years into a technology is happy when there is an uptake
09:05
<mhausenblas>
saying this is not the case is a simply dishonest
09:05
<othermaciej>
I like it when people use WebKit
09:05
<mhausenblas>
s/is a/is
09:05
<Philip`>
I guess it still could be useful for each group to point out the unjustified fanboyism of each other group, even if we're all doing it
09:05
<mhausenblas>
Philip`: +1
09:06
<othermaciej>
but I have distinctly mixed feelings if they use it in a way where they make significant changes which they don't contribute back
09:06
<othermaciej>
I think what's most important is to be aware of your own unjusfified fanboyism
09:06
<othermaciej>
and try to keep it from influencing technical decisions in a bad way
09:06
<Philip`>
It seems the kind of situation that makes self-awareness difficult, which is why it's helpful to have other people point it out
09:09
<Philip`>
(...though preferably in a way that's not too rude or condescending, I guess)
09:09
<mhausenblas>
as for me I just say: whatever data comes in there I'm happy (bonus if it links to other data, as this is the Web, right?)
09:09
<mhausenblas>
two examples: just started reading the wonderful RESTful Web Services book - awesome. so much useful stuff in it (and I see perfectly how linked data connects to it)
09:09
<mhausenblas>
other example: started to work with Erik on RESTful RDF (http://dret.typepad.com/dretblog/2009/05/rest-and-rdf-granularity.html)
09:10
<mhausenblas>
Philip`: yes. being polite and honest doesn't exclude each other ;)
09:12
<mhausenblas>
ok, back to some hacking ... cya laters
09:23
<othermaciej>
"These are the kind of pipe dreams that I used to ridicule semantic web folk about back before they body-snatched me!"
09:23
<othermaciej>
now that's a high degree of self-awareness
10:03
<mhausenblas>
(just FYI, the discussion here earlier motivated me to do a quick write-up for starting to collect RDF anti-patterns)
10:03
<mhausenblas>
see http://chatlogs.planetrdf.com/swig/2009-05-24.html#T09-00-52
10:03
<mhausenblas>
(back to hacking now, really ;)
11:52
<annevk42>
othermaciej, am now
11:53
<othermaciej>
annevk42: I was going to ask you to remind me how to check out and generate the Design Principles draft
11:53
<othermaciej>
I don't think I have a checkout of the spec here
11:57
<annevk42>
http://dev.w3.org/cvsweb/
11:58
<annevk42>
it's in html5/html-design-principles
11:58
<annevk42>
you want to edit Overview.src.html
11:59
<annevk42>
mhausenblas, nofollow is defined in the HTML5 spec itself
11:59
<annevk42>
mhausenblas, so it doesn't need to be on RelExtensions
11:59
<mhausenblas>
ok
12:07
<othermaciej>
what's the right place to send EventSource feedback?
12:07
<annevk42>
public-webapps⊙wo
12:28
<othermaciej>
hmm I can't seem to remember my w3c cvs password
12:29
<othermaciej>
guess I'll have to ask a cvs admin for help
12:32
<annevk42>
it works with an ssh keypair or whatever that's called
13:48
gsnedders
wonders why the wifi has been so slow
13:48
<gsnedders>
Suddenly it jumped just as I said that from 10 to 200 KB/s
13:48
<gsnedders>
And back down
13:49
gsnedders
is trying to download the spec :P
13:53
<hsivonen>
whoa Anolis adds more bytes than the split-out sections take away
13:54
<gsnedders>
Back down to a number of bytes per second
14:28
<hsivonen>
Hixie: increased the limit to 4 MB
14:41
<karlcow>
[04:06] <othermaciej> it does seem like there is some degree of fanboyism around RDFa
14:41
<karlcow>
s/RDFa/whatwg|html5|hixie|$fan/ ($fan = pick your own fav topics of a community), which makes the sentence stupid
14:41
<karlcow>
[04:08] <Philip`> It seems there's some degree of fanboyism around every technology
14:41
<karlcow>
fully agreed.
15:10
<Philip`>
karlcow: It doesn't make the sentence stupid, it just makes it less general than it perhaps should be
15:19
gsnedders
guesses he ought to do school work again
15:19
gsnedders
sighs
15:25
<annevk4>
/whois Hixie
15:26
annevk4
tries to figure out why HTTP is no longer functioning
15:31
gsnedders
wonders what IRI to use for atom:category@scheme in a spec
15:32
<zcorpan_>
... he refers to a note from Simon Pieters
15:32
<zcorpan_>
... I am sure we addressed that."
15:33
<zcorpan_>
i have not received a reply to the email that björn cited
16:10
<gsnedders>
Nice. Back down to 50% test cases failing.
16:44
jgraham
wonders if he has missed anything, sees a depressing amount of email
17:11
gsnedders
is thankful html5lib has so many test cases with how much he's broken it
17:18
<jgraham>
gsnedders: Did anyone ever tell you that the order of the errors isn't significant in the tokenizer tests
17:18
<gsnedders>
jgraham: No
17:18
<jgraham>
Oh well they should have done
17:18
<jgraham>
At least I think that is true
17:18
<gsnedders>
This doesn't account for 500 test cases failing, though :D
17:19
<jgraham>
No. But It might account for some ailing sonce you fix the major issues
17:19
<gsnedders>
I think it fixed one, the order of errors.
17:25
<jgraham>
Philip`: I plan to implement microdata-to-*
17:26
gsnedders
gets down to 61 errors
17:27
<gsnedders>
Huh. Somehow it goes from afterAttributeValueQuoted to attributeName
17:29
<gsnedders>
Oh, wait. I hadn't re-run the tests since I fixed that illogic.
17:32
<gsnedders>
And down to 25…
17:49
<Philip`>
jgraham: I told him that the order of errors isn't significant when the test case has ignoreErrorOrder:true or whatever it is
17:56
<gsnedders>
Which is different to what jgraham asked.
18:01
<Philip`>
gsnedders: Indeed
18:01
<Philip`>
The idea is that the tests which don't have ignoreErrorOrder ought to have predictable error order in any sane implementation, so you should test the order in order to detect bugs, though technically it's not required by the spec
18:06
<gsnedders>
2000 line diff. Fun.
18:07
jgraham
notes that a PHP implementation might not be considered sane :)
18:08
gsnedders
agrees
18:09
gsnedders
has been adding things to his post entitled "PHP Grievances" while working on html5lib over the past few days :)
18:14
<gsnedders>
Philip`: Oh, and re: yesterday, there is look-behind in the data-state.
18:17
<takkaria>
gsnedders: there doesn't have to be
18:17
<gsnedders>
There isn't in Python or PHP implementations, dunno about Ruby…
18:17
<gsnedders>
takkaria: And seeming I just changed the PHP impl. to be like that… :D
18:17
<gsnedders>
http://code.google.com/p/html5lib/source/detail?r=2aa8163f5b86662263bd609f81e40c324d574e42
18:18
<takkaria>
I wonder what we did in hubbub for that
18:20
<takkaria>
oh yeah
18:21
<takkaria>
basically, hubbub uses an inputstream which doesn't support ungetting or peeking backward
18:22
<gsnedders>
Ow.
18:22
<takkaria>
acutally it's a highly efficient implementation
18:22
<takkaria>
we have a current position in the inputstream and then we have an offset number of bytes we are into the buffer
18:23
<gsnedders>
Oh, sure. But it's not going to be fun to implement parts of the spec with :)
18:23
<takkaria>
indeed not
18:23
<takkaria>
but it's not much pain really
18:35
<gsnedders>
So splitting out the input stream was around a 0.1s hit when tokenizing the spec
18:35
<gsnedders>
(Now around 6.1s)
19:01
<takkaria>
that's not terrible
19:04
<Philip`>
gsnedders: Oh, I forgot about that bit
19:04
<Philip`>
but it's only four characters, and only needed in rare circumstances, so it's not much of a problem to add a special look-behind buffer
19:05
Philip`
supposes it'd be possible to optimise it to a single int keeping track of the state, instead of keeping all the characters
19:21
<takkaria>
it largely depends on how performant string manipulation is in PHP
19:23
hsivonen
notes that the tokenizer can be implemented without lookagead or lookbehind
19:23
<hsivonen>
*lookahead
19:23
<takkaria>
at some point I should really look at the v.nu parser and see how you remove lookahead
19:24
<takkaria>
because I can imagine it being a fairly decent performance benefit
19:55
<gsnedders>
takkaria: Basically, I think the biggest limit when dealing with the data state is going to be function call overhead.
19:55
<gsnedders>
takkaria: Function call overhead is a non-negilable amount of time when tokenizing the spec, so basically have to rely as much as possible on language constructs and not functions.
19:57
<Philip`>
Write the whole parser as a single giant regular expression, then you won't have any non-native function calls
20:02
<gsnedders>
Philip`: "The maximum length of a compiled pattern is 65539 (sic) bytes if PCRE
20:02
<gsnedders>
is compiled with the default internal linkage size of 2."
20:02
<gsnedders>
Philip`: Also: "All values in repeating quantifiers must be less than 65536."
20:06
<takkaria>
gsnedders: I'm sure that the function call overhead is a pretty big portion, but string use will be a factor
20:06
<takkaria>
hubbub tries to delay string handling to as late as possible by using counters and indexes into buffers instead of doing string ops
20:06
<takkaria>
and then does the string ops when tokens are emitted
20:07
<takkaria>
I dunno, it just might be an avenue you could look at if you're intending on optimising further
20:07
<gsnedders>
http://stuff.gsnedders.com/phphtml5lib.pdf
20:07
<gsnedders>
takkaria: We work one byte at a time
20:07
<gsnedders>
takkaria: (Basically, we use a single byte long string)
20:07
<Philip`>
Do you have anything like charsUntil?
20:08
<Philip`>
(from Python's html5lib)
20:08
<gsnedders>
Philip`: Yes
20:08
<Philip`>
Does that work one byte at a time?
20:08
<gsnedders>
No
20:08
<gsnedders>
Well, I guess within PHP it does :P
20:08
<gsnedders>
(Actually, I think PHP just uses libc for strspn, so within libc it does)
20:09
<Philip`>
Okay, so when you say you work one byte at a time I should ignore you :-)
20:09
<gsnedders>
Look at that graph: the big drop was when I added that.
20:09
<gsnedders>
Apart from when we grab a lot of stuff at once, we work one byte at a time :D
22:05
<Philip`>
Does SGML precisely define some concept like case-insensitivity?
23:02
gsnedders
pulls out his battered second hand copy of the SGML Handbook
23:04
<Philip`>
Hmm, I prefer my books to be covered with breadcrumbs
23:05
<gsnedders>
I think not
23:06
<gsnedders>
No, it does
23:06
<gsnedders>
Uppercase letters (A-Z, numbers 65–90) map to their lowercase equivalents (a–z, numbers 97–122).
23:07
<gsnedders>
You have to look in about five different places to find that, but it seems well enough defined.
23:07
<Philip`>
Ah, thanks
23:11
<gsnedders>
Hmm, I have an unkillable Terminal
23:12
<gsnedders>
Like, running kill -KILL against it has no effect
23:15
<Philip`>
That's easy to fix - reboot
23:15
<gsnedders>
That means terminating everything else, though
23:16
<Dashiva>
How do you know it's unkillable? E.g. can you still use it?
23:16
<gsnedders>
Dashiva: No
23:16
<gsnedders>
Oh well, reboot it is.
23:17
<Philip`>
Can't you just ignore it?
23:17
<gsnedders>
I can't open stuff now either. WTF?
23:17
<gsnedders>
brb
23:24
<gsnedders>
wee… sanity returns!
23:26
<Dashiva>
You might want to question the sanity of your computer if a stubborn terminal can prevent programs from opening
23:27
<gsnedders>
It's my computer, so of course it fails to meet the definition of sane.
23:32
Philip`
is reminded of the fun times when his gamin server dies, and so any program that tries to connect to it on startup hangs, and eventually when he realises the problem and restarts the daemon suddenly dozens of copies of Krusader and KPDF spring into existence as they get unblocked