MCPcopy Create free account
hub / github.com/dask/dask / test_describe

Function test_describe

dask/dataframe/tests/test_dataframe.py:407–474  ·  view source on GitHub ↗
(include, exclude, percentiles, subset)

Source from the content-addressed store, hash-verified

405 ],
406)
407def test_describe(include, exclude, percentiles, subset):
408 data = {
409 "a": ["aaa", "bbb", "bbb", None, None, "zzz"] * 2,
410 "c": [None, 0, 1, 2, 3, 4] * 2,
411 "d": [None, 0, 1] * 4,
412 "e": [
413 pd.Timestamp("2017-05-09 00:00:00.006000"),
414 pd.Timestamp("2017-05-09 00:00:00.006000"),
415 pd.Timestamp("2017-05-09 07:56:23.858694"),
416 pd.Timestamp("2017-05-09 05:59:58.938999"),
417 None,
418 None,
419 ]
420 * 2,
421 "f": [
422 np.timedelta64(3, "D"),
423 np.timedelta64(1, "D"),
424 None,
425 None,
426 np.timedelta64(3, "D"),
427 np.timedelta64(1, "D"),
428 ]
429 * 2,
430 "g": [True, False, True] * 4,
431 }
432
433 # Arrange
434 df = pd.DataFrame(data)
435 df["a"] = df["a"].astype(get_string_dtype())
436
437 if subset is not None:
438 df = df.loc[:, subset]
439
440 ddf = dd.from_pandas(df, 2)
441
442 # Act
443 actual = ddf.describe(
444 include=include,
445 exclude=exclude,
446 percentiles=percentiles,
447 )
448 expected = df.describe(
449 include=include,
450 exclude=exclude,
451 percentiles=percentiles,
452 )
453
454 if "e" in expected:
455 expected = _drop_mean(expected, "e")
456
457 assert_eq(actual, expected)
458
459 if PANDAS_GE_310:
460 # The 'include' and 'exclude' arguments are deprecated for Series.describe and
461 # will be removed in a future version. These arguments have no effect on Series
462 # and will be removed.
463 series_kwargs = {}
464 else:

Callers

nothing calls this directly

Calls 6

describeMethod · 0.95
get_string_dtypeFunction · 0.90
assert_eqFunction · 0.90
_drop_meanFunction · 0.70
astypeMethod · 0.45
describeMethod · 0.45

Tested by

no test coverage detected