Home > mailing lists

Re: ANALYZE sampling is too good - Mailing list pgsql-hackers

From	Greg Stark
Subject	Re: ANALYZE sampling is too good
Date	December 6, 2013 19:06:31
Msg-id	CAM-w4HPDaioC9epxviuNkD-8ZnYeBSb7Z=uHQUkTohMfdkgFVQ@mail.gmail.com Whole thread Raw
In response to	Re: ANALYZE sampling is too good (Andres Freund <andres@2ndquadrant.com>)
Responses	Re: ANALYZE sampling is too good
List	pgsql-hackers

Tree view

It looks like this is a fairly well understood problem because in the
real world it's also often cheaper to speak to people in a small
geographic area or time interval too. These wikipedia pages sound
interesting and have some external references:

http://en.wikipedia.org/wiki/Cluster_sampling
http://en.wikipedia.org/wiki/Multistage_sampling

I suspect the hard part will be characterising the nature of the
non-uniformity in the sample generated by taking a whole block. Some
of it may come from how the rows were loaded (e.g. older rows were
loaded by pg_restore but newer rows were inserted retail) or from the
way Postgres works (e.g. hotter rows are on blocks with fewer rows in
them and colder rows are more densely packed).

I've felt for a long time that Postgres would make an excellent test
bed for some aspiring statistics research group.


-- 
greg

pgsql-hackers by date:

From: Tom Lane
Date: 06 December 2013, 19:03:01
Subject: Re: Proof of concept: standalone backend with full FE/BE protocol

From: Tom Lane
Date: 06 December 2013, 19:10:31
Subject: Re: pg_archivecleanup bug

Re: ANALYZE sampling is too good - Mailing list pgsql-hackers

Previous

Next