Home > mailing lists

Re: Predicting query runtime - Mailing list pgsql-general

From	Istvan Soos
Subject	Re: Predicting query runtime
Date	September 13, 2016 19:54:40
Msg-id	CALdQGguuuVm0-WtDKy+_pDjghXbRxEWULAbn-t+26q2VTkvSPw@mail.gmail.com Whole thread
In response to	Re: Predicting query runtime (Vinicius Segalin <vinisegalin@gmail.com>)
Responses	Re: Predicting query runtime
List	pgsql-general

Tree view

On Tue, Sep 13, 2016 at 2:06 AM, Vinicius Segalin <vinisegalin@gmail.com> wrote:
> 2016-09-12 18:22 GMT-03:00 Istvan Soos <istvan.soos@gmail.com>:
>> At Heap we have non-trivial complexity in our analytical queries, and
>> some of them can take a long time to complete. We did analyze features
>> like the query planner's output, our query properties (type,
>> parameters, complexity) and tried to automatically identify factors
>> that contribute the most into the total query time. It turns out that
>> you don't need to use machine learning for the basics, but at this
>> point we were not aiming for predictions yet.
>
> And how did you do that? Manually analyzing some queries?

In this case, it was automatic analysis and feature discovery. We were
generating features out of our query parameters, out of the SQL
string, and also out of the explain analyze output. For each of these
features, we have examined the P(query is slow | feature is present),
and measured its statistical properties (precision, recall,
correlations...).

With these we have built a decision tree-based partitioning, where our
feature-predicates divided the queries into subsets. Such a tree could
be used for predictions, or if we would like to be fancy, we could use
the feature vectors to train a neural network.

Hope this helps for now,
  Istvan

pgsql-general by date:

From: Steve Crawford
Date: 13 September 2016, 18:23:30
Subject: Re: Installing 9.6 RC on Ubuntu [Solved]

From: Oleg Bartunov
Date: 13 September 2016, 20:12:48
Subject: Re: Predicting query runtime

Re: Predicting query runtime - Mailing list pgsql-general

Previous

Next