Home


python notes

See also:

~/notes/pip.md
~/notes/pyenv.md
~/notes/pynamo.md
~/notes/numpy.md
~/notes/pandas.md

Gotcha when slicing with -1

range(3, -1, -1) and slice(3, -1, -1) are very different.

You can range(3, -1, -1) to count down from 3 to zero (inclusive), but if you slice s[3:-1:-1] you always get an empty string.

In slices, -1 always means the last element in the list, regardless of the step.

s = 'abcdefg'  
s[3 :  2 : -1]  # => 'd'  
s[3 :  1 : -1]  # => 'dc'  
s[3 :  0 : -1]  # => 'dcb'  
s[3 : -1 : -1]  # => ''  

Remember that -1 in start or stop of s[start:stop:step] means the last element.

And remember to put the slice args in the right direction when step is -1

"abcdef"[3:1:-1] # => 'dc' (first arg is inclusive, second is exclusive)  

Walrus major gotcha coming from ruby

This pattern works great in ruby to do something if the return is not nil

if val = vend_value_or_nil  
  # Do something with val  
end  

but in python, do NOT do this if the return can be 0, empty string, or an empty container:

if val := vend_value_or_none:  
    # Bad! value can be 0 which is falsy  

the safer version is not nearly as elegant, so maybe use another pattern entirely:

if (val := vend_value_or_none) is not None:  
    # Do something with val  

use:

val = vend_value_or_none()  
if val is not None:  
    # Do something with val  

Use min with a key instead of insane code

If I find myself ever doing something awful like this:

min_val = float("inf")  
min_key = -1  
for k, v in my_counter.items():  
    if v < min_val:  
        min_val = v  
        min_key = k  

use this instead:

min_key, min_val = min(my_counter.items(), key=lambda item: item[1])  

defaultdict gotcha

Use defaultdict(int) not defaultdict(0).

If I need a number other than 0 to start with, use defaultdict(lambda: 5)

For the 0 case, prob better to use from collections import Counter

Use for/else

If I find myself creating a bit of state called found and then breaking from a loop and testing:

found = False  
for ...:  
    if meets_found_condition:  
        found = True  
        break  

if not found:  
    do_default  

Use this instead:

for ...:  
    if meets_found_condition:  
        break  
else:  
    do_default  
      

The else runs when the loop finishes without break, including when the loop has zero iterations.

Instance of, instanceof

x = 1  
type(x) # => int  
isinstance(x, (int, float)) # => True  

class C:  
    pass  
class D(C):  
    pass  

isinstance(D(), C) # => True  
type(D()) is C     # => False  

Unix timestamp

from time import time  
time()  

See also ~/dev/snippets/python/date_time.py
~/dev/snippets/python/get_timezone.py
~/dev/snippets/python/hours_between_two_datetimes.py
~/dev/snippets/python/time_execution.py

NamedTuple new interface

Instead of

from collections import namedtuple  
Point = namedtuple("Point", ["x", "y"])  

I can use

from typing import NamedTuple  

class Point(NamedTuple):  
    x: int  
    y: int  
Resize = namedtuple("Resize", ["width", "height"])  

Deserializing into a structured type

Nicer than:

points = json.loads('{"points": [{"x": 0, "y": 0}]}')  

is decoding into structured objects.

First option, namedtuple:

import json  
from collections import namedtuple  

Point = namedtuple("Point", ["x", "y"])  
data = json.loads('{"points": [{"x": 0, "y": 0}]}')  
points = [Point(**p) for p in data["points"]]  

Second option, dataclasses. mutable by default but can use frozen=True for immutability, or use eq=False to put them in sets:

import json  
from dataclasses import dataclass  

@dataclass  
class Point:  
    x: int  
    y: int  

data = json.loads('{"points": [{"x": 0, "y": 0}]}')  
points = [Point(**p) for p in data["points"]]  

Third option, pydantic, closest to Swift’s Decodable:

from pydantic import BaseModel  

class Point(BaseModel):  
    x: int  
    y: int  

class Payload(BaseModel):  
    points: list[Point]  

payload = Payload.model_validate_json(  
    '{"points": [{"x": 0, "y": 0}]}'  
)  

If I already have a dictionary, use Payload.model_validate(data) instead.

Pointer equality

Use is, it’s analogous to swift’s ===

Multiline string python

If I want newlines:

text = """First line  
Second line  
Third line"""  

If I don’t want newlines:

text = (  
    "This is a long sentence "  
    "continued on another source line."  
)  

If I want newlines but no indents:

from textwrap import dedent  

dedent("""\  
    First line  
    Second line  
    Third line  
""")  

Trim whitespace

x.strip()  
x.lstrip()  
x.rstrip()  

Or eliminate internal spaces

x.replace(' ', '')  

Tip to remove whitespace and indents

from textwrap import dedent  

text = dedent("""  
    Hello  
    world  
""").strip()  
# 'Hello\nworld'  

Tip for printing vars

yes:

  long_name = 1  
  print(f'{long_name=}')  

no:

  long_name = 1  
  print(f'long_name={long_name}')  

pdb tips for moving around

For vanilla pdb:

(Pdb) w        # Show the call stack  
(Pdb) u        # Select the caller’s frame  
(Pdb) list     # Show the code around the caller  
(Pdb) p node   # Inspect that caller’s variables  
(Pdb) d        # Move back toward the current frame  

python gotcha with dataclasses

Using a dataclass makes types not hashable

@dataclass  
class Node:  
    value: int  

{Node(1)} # => runtime error  

But it does gives nodes field level equality

Node(1) == Node(1) # => true  

I can use

@dataclass(eq=False)  

or

@dataclass(frozen=True)  

or implement __hash__ or __eq__ myself.

Frozen is best for immutable values like points that should compare equal (e.g.
(1,2) == (1,2) makes sense to be true). But frozen is not best for nodes,
because Node(1) and Node(1) I likely do not want to be equal. I want identity instead (use eq=False)

How to get interactive help in a repl

  1. Append a ? to the fn or method I’m trying to understand

    open?

  2. Append ?? to see docstring and source!

    import requests
    requests??

  3. Use the help built-in (I prefer ??):

    help(sorted)

  4. Print docstring directly:

    print(open.__doc__)

  5. Use dir, the old standby:

    x = “abc”
    dir(x)

How to memoize

Before 3.9:

from functools import lru_cache  
@lru_cache(maxsize=None)  
def fn...  

After 3.9:

from functools import cache  
@cache  
def fn...  

Timing functions

If in ipython or jupyter notebook, use magic %timeit

For a single pass, use perf_counter, which is guaranteed to be monotonic:

import time  

t = time.perf_counter()  
# do work  
print(time.perf_counter() - t)  

For repeated runs and summary statistics, use timeit:

import timeit, statistics  

def work():  
    pass  

times = timeit.repeat(work, number=1000, repeat=5)  

print("runs:", times)  
print("min:", min(times))  
print("max:", max(times))  
print("mean:", statistics.mean(times))  
print("stdev:", statistics.stdev(times))  

Can get a ruby-ish feel with a context manager:

import time  
from contextlib import contextmanager  

@contextmanager  
def timed_block(label: str = "timed_block"):  
    start = time.perf_counter()  
    try:  
        yield  
    finally:  
        end = time.perf_counter()  
        print(f"{label}: {(end - start) * 1000:.3f} ms")  

with timed_block():  
    total = sum(i * i for i in range(10_000_000))  

context manager basic usage

Using the decorator:

from contextlib import contextmanager  

@contextmanager  
def example():  
    print("enter")  
    try:  
        yield "some value"  
    finally:  
        print("exit")  

with example() as value:  
    print(value)  

Or manually creating a class and implementing the two required dunder methods:

class Example:  
    def __enter__(self):  
        print("enter")  
        return "some value"  

    def __exit__(self, exc_type, exc_value, traceback):  
        print("exit")  
        return False  # Do not suppress exceptions raised raised inside the with block  

with Example() as value:  
    print(value)  

context manager mistake

It is valid to return from inside a with statement.

Do not do this:

unsafe_lock = Lock()  
unsafe_lock.acquire_lock()  
res = n + 1  
unsafe_lock.release()  
return res  

Just do:

with unsafe_lock:  
    return n + 1  

Get multiple indices of a list

from operator import itemgetter  
x = ['a', 'b', 'c']  
indices = [0, 2]  
itemgetter(*indices)(x) # => ('a', 'c')  

Or, if I can’t remember all that

x = ['a', 'b', 'c']  
indices = [0, 2]  
[x[i] for i in indices]  

Or

x = ['a', 'b', 'c']  
[y for i,y in enumerate(x) if i in [0,2]]  

Use StrEnum if I need rawvalues of strings

from enum import StrEnum  

class Role(StrEnum):  
    SYSTEM = "system"  
    USER = "user"  

Role.SYSTEM.value # => "system"  

Use auto to skip labeling enums

from enum import Enum, auto  

class Role(Enum):  
    SYSTEM = auto()  
    USER = auto()  

Use classmethod decorator for class methods

class C:  
    @classmethod  
    def y(cls):  
        return cls  

C is C.y() # => True  

Use array for homogenous primitives

import array  
array.array('I', [1,2])  

How to use sleep

from time import sleep  
sleep(0.1)  

asyncio run main

Basic setup to get an async context in python:

import asyncio  

async def main() -> None:  
    print("sleeping for one second...")  
    await asyncio.sleep(1)  
    print("done!")  

asyncio.run(main())  

Pydantic basic usage

Pydantic doesn’t use positional arguments:

from pydantic import BaseModel  

class User(BaseModel):  
    name: str  

# Create a user from safe input  
user = User(name="a")  

# Or  
user_data = {'name': 'a'}  
user = User(**user_data)  

# Create a user from unknown user input  
User.model_validate({'name': 1})  

# Or   
User.model_validate_json('{"name": "lou"}')  

# Get json:  
user.model_dump_json()  

Python’s GC

Primary mechanism is ref counting:

In CPython, the primary algorithm for garbage collection is reference  
counting. Essentially, each object keeps count of how many references point  
to it. As soon as that refcount reaches zero, the object is immediately  
destroyed: CPython calls the __del__ method on the object (if defined) and  
then frees the memory allocated to the object. In CPython 2.0, a  
generational garbage collection algorithm was added to detect groups of  
objects involved in reference cycles  

Fluent Python pg 219

tuples are immutable

x = (1,)  
id(x)  
x += (2,)  
id(x) # => different ID  

But be careful when a tuple holds a mutable instance:

x = ["a"]  
y = (x,)  
x.append("b")  
y  
(['a', 'b'],)  

types with value semantics

Get bytecode

def fn():  
    return "hello".count("l")  

Raw bytecode:

fn.__code__.co_code  

Readable bytecode:

import dis  
dis.dis(fn)  

Or:

list(dis.get_instructions(fn))  

ruff linter can fix in place

I’ve been using flake8, but starting to use ruff instead. E.g.

ruff check  
ruff format  

assert is a statement

If I use parens, it’s just grouping, not a function call:

assert(x == 1)  

Is actually

assert (x == 1)  

It’s as if I were writing:

if(true):  

Which I would not do.

Python visualizer

https://pythontutor.com/render.html#mode=display

Disallow named params

Use the /

def fn(a, /):  
    pass  

fn(a="1")  # Raises TypeError  

Remove from dictionary without raising

d = {'a': 'b'}  
d.pop('c', None)  

Iterate over a dict, python gotcha

This does not do what I expect:

for k, v in { 'ab': 12 }:  
    print(k)  
    print(v)  
# => prints 'a' followed by 'b', because I'm actually unpacking the key only  

To destructure the keys and values from a tuple at each iteration, use items():

for k, v in {'ab': 12 }.items():  
    print(f'k: {k}, v: {v}')  

Pydantic immutable record

I always liked NameTuple. Pydantic went with mutable by default, which makes
sense for the domain. I can use this where I want immutability:

from pydantic import BaseModel, ConfigDict  
  
class User(BaseModel):  
    id: int  
    name: str  
    model_config = ConfigDict(frozen=True)  

Then:

 u = User(id=1, name="foo")  
 u.id=2  
 # => Raises a ValidationError  

Type hints in 3.9+

Importing typing is no longer necessary.
But, I still need mypy as the static checker.
There is nothing built into python to static check your types.

Type hints refresher

a: tuple[int, str] = (1, "hi")  
b: list[str] = ["a", "b"]  
c: dict[str, str] = {"a": "b"}  
d: set[int] = {1, 2, 3}  

from collections.abc import Callable  
e: Callable[[int], int] = lambda x: x + 1  
f: Callable[[int, int], int] = lambda x, y: x + y  
g: Callable[[int | None], int] = lambda x: x if x else 0  

For optionals, I can either use:

#  3.10+ only  
h: int | None = None  
i: int | None = 1  

or:

from typing import Optional  
j: Optional[int] = None  
k: Optional[int] = 1  

or:

from typing import Union  
l: Union[int, None] = None  
m: Union[int, None] = 1  

Set refresher

s = set()  
s.add(1)  
1 in s # => True  

NamedTuple refresher

NamedTuples are immutable:

from typing import NamedTuple  
class Point(NamedTuple):  
    x: int  
    y: int  
p = Point(x=1, y=0)  
p.y = 1  
# => Raises AttributeError  

You can put these in a set:

{ Point(0,0) }  

or dictionary, and keying works how I’d expect:

d = { Point(0,0): 'a'}  
x = Point(0,0)  
y = Point(0,1)  
d[x] # => 'a'  
d[y] # => KeyError  

Dataclass refresher

Dataclasses are mutable:

from dataclasses import dataclass  

@dataclass  
class Point:  
    x: int  
    y: int  
p = Point(x=1, y=0)  
p.y = 1  
# => Does not throw  

Frozen dataclasses and a gotcha

from dataclasses import dataclass  

@dataclass(frozen=True)  
class D:  
    field = 1  

d = D()  
d.field = 2 # => raises FrozenInstanceError, as expected  
d.field # => 1  
D.field = 99  
d.field # => 99  

Dataclasses with field factory for list

Don’t do this:

@dataclass  
class D:  
    l = []  

D().l.append('a')  
D().l  
# => ['a']  

Do this:

from dataclasses import dataclass, field  
@dataclass  
class D:  
    l: list[int] = field(default_factory=list)  

D().l.append('a')  
D().l  
# => []  

dataclasses can have a default field value

from __future__ import annotations  
from dataclasses import field  

@dataclass(eq=False)  
class Node:  
    val: int  
    neighbors: list[Node] = field(default_factory=list)  

dataclasses can use post init to set default values

__post_init__ is a dataclass hook, it is not part of the standard object model.

from dataclasses import dataclass, field  

@dataclass  
class C:  
    quota: int  
    remaining: int = field(init=False)  

    def __post_init__(self):  
        self.remaining = self.quota  

     def use(self, amount: int):  
        self.remaining -= amount  

dataclasses field reference

https://docs.python.org/3/library/dataclasses.html#dataclasses.field

Immutable dict

It’s not hashable, but it does guarantee elements can’t chance:

from types import MappingProxyType  

x = MappingProxyType({'a': 1})  
x['a'] = 2 # => raises 'TypeError: 'mappingproxy' object does not support item assignment'  

Note a mutable type still allows mutation in the collection:

x = MappingProxyType({'a': [1]})  
x['a'].append(2) # Fine  
x['a'] # => [1, 2]  

Ternary gotcha

I always assume this is valid python, it’s not: x ? x : 0
I want: x if x else 0

Check MRO on a type

Foo.__mro__  

Debugging

pdbpp is an absolute must. I don’t use the following any more. I just stick to pdb and pdbpp with this in ~/.pdbrc.py

from pdb import DefaultConfig  

class Config(DefaultConfig):  
    sticky_by_default = True  

No longer used, because ipdb doesn’t play super nicely with pdbpp:

`ipdb` is also handy, because I can get it show additional lines around a breakpoint.  
  
Add this to `~/.bash_profile`  
  
    # Use ipdb when python hits a `breakpoint()`  
    export PYTHONBREAKPOINT="ipdb.set_trace"  
  
    # Show additional context around ipdb breakpoints  
    # Change this by typing 'context NUM' when a breakpoint is hit. Then 'bt' to redraw the source.  
    export IPDB_CONTEXT_SIZE=5  

With that config, I don't have to explicitly breakpoint with `import ipdb; ipdb.set_trace(context=5)`.  
I can use `breakpoint()` and then either `sticky` if I want to use pdbpp or `context NUM` followed by `bt` if I just want to see more.  

Start repl with async support

python -m asyncio  

Install a modern version on AL2023

https://github.com/amazonlinux/amazon-linux-2023/issues/483#issuecomment-1928446605
^ Don’t do this. Do this:

sudo dnf install python3.12  

Walrus

Available in python 3.8+

if (n := "world"):  
    print(f"{n} hello")  

Define all public exports

I prefer explicit imports, but I see this often. __all__ defines what should be imported on import * from xyz

__all__ = ["MyClass"]  

class MyClass():  
    pass  

Defaultdict usage

from collections import defaultdict  
x = defaultdict(str)  
x['y'] # => ''  

Or commonly with a list

x = defaultdict(list)  
x['y'].append(1)  

sorted always returns a list

sorted(['b', 'a']) # => ['a', 'b']  
sorted('ba') #=> ['a', 'b']  
sorted(('b', 'a')) # => ['a', 'b']  

Lists are not hashable

Only immutable objects can be hashed and added to sets / dicts.
The trick is to use a tuple instead:

s = set()  
x = ['a', 'b']  
s.add(x) # => raises TypeError  
s.add(tuple(x))  

Split a string into list

"a b".split(" ") # => ["a", "b"]  
list("ab") # => ["a", "b"]  

Join a list into str

''.join(['a', 'b'])  

Modern typing for callable

Old:

from typing import Callable  

New:

from collections.abc import Callable  

The Callable syntax is Callable[[arg1, arg2], ret].

To support arbitrary args, use Callable[..., ret].

Protocol usage

from typing import Protocol  
from typing import TypeVar  

class ExampleProtocol(Protocol):  
    """  
    A common interface for types that provide a `foo` property, which  
    can be helpful to satisfy mypy when duck typing / using generics.  
    """  
    @property  
    def foo(self) -> float:  
        ...  


# Later  
# from whatever.example_protocol import ExampleProtocol  
T = TypeVar('T', bound=ExampleProtocol)  

def requires_a_foo(x: T):  
    print(x.foo)  

Lists can be assigned by slice

x = [0, 1, 2, 3]  
s = slice(None, None, 2)  
x[s] = ['a', 'b']  
x  
# =>  ['a', 1, 'b', 3]  

Using a key with sorted

pairs = [(1, "Charlie"), (2, "Alice"), (3, "Bob")]  
sorted(pairs, key=lambda pair: pair[1])  
# [(2, "Alice"), (3, "Bob"), (1, "Charlie")]  

Mistake: Use list comprehension instead of filter

nums = [1, 2, 3, 4, 5]  

Instead of:

list(filter(lambda n: n % 2 == 0, nums))  

Use:

[n for n in nums if n % 2 == 0]  

Mistake: don’t index into a list of tuples, prefer unpacking

samples = [(1, ['a']), (2, ['a', 'b'])]  

Instead of:

for sample in samples:  
    stack = sample[1]  
    idx = sample[0]  
    pass  

Use:

for idx, stack in samples:  
    pass  

It’s cool that that works.

Mistake: use more elegance for membership

I was writing a sample tokenizer and did this:

elif char in ['+', '*', '(', ')']:  

better to write:

elif char in '+*()':  

Mistake: use more elegance for comparing against None

No:

x != None  

Yes:

x is not None  

python 3.12 added ‘batched’ to itertools

[list(x) for x in batched('ABCDEFG', 3)]  
[['A', 'B', 'C'], ['D', 'E', 'F'], ['G']]  

I always think I want something called grouped or chunked, but it’s batched!

I sometimes also confuse batched with divide (up next).

There is no ‘divide into x groups’ in itertools

from more_itertools import divide  
divide(3, 'ABCDEFG')  
  

Asterisk in the arg list

The bare * means all parameters after it must be passed by keyword:

def grouper(iterable, n, *, incomplete='fill', fillvalue=None)  

Example from itertools recipes: docs.python.org/3/library/itertools.html

Get the next element from an iterator

it = iter(range(10))  
next(it) # => 0  

Get the first elements from an iterator or generator

from itertools import islice  
gen = (x * x for x in range(1, 10))  
list(islice(gen, 0, 3))  
# => [1, 4, 9]  

takewhile

list(takewhile(lambda x: x > 2, [4,3,2,1]))  

zip

zip stops at the shortest arg, but there is a flavor in itertools:

from itertools import zip_longest  
list(zip_longest('abc', [1,2], fillvalue=-1))  
# => [('a', 1), ('b', 2), ('c', -1)]  

Get first elements from generator

import itertools
first5 = itertools.islice(gen, 5)
print(list(first5))

Decode bytes to utf8 string

my_bytes.decode('utf8')  

try except bare except

  1. Python does not automatically bind error.

  2. Catch specific exceptions when possible:

     try:  
         do_work()  
     except MyError as err:  
         print(err)  
  3. Use except Exception: for a broad catch of ordinary errors.
    Avoid bare except: because it also catches KeyboardInterrupt and SystemExit.

The idiomatic way to do something like swift’s error cases:

enum PipelineError: Error {  
    case invalidInput  
    case unsupportedFormat(String)  
}  

Is to subclass Exception:

class PipelineError(Exception):  
    pass  

class InvalidInputError(PipelineError):  
    pass  

class UnsupportedFormatError(PipelineError):  
    pass  

then:

raise UnsupportedFormatError("Unsupported format: TIFF")  

and handle with:

try:  
    run_pipeline()  
except InvalidInputError:  
    print("Check the input")  
except UnsupportedFormatError as err:  
    print(err)  
except PipelineError:  
    print("Some other pipeline error")  

View the MRO with:

UnsupportedFormatError.mro()  

It’s ok to define a var in a catch block

# Do not need to define x here  
try:  
    raise ValueError("broken precondition")  
except:  
    x = 1  
x += 1  # => 2  

Huh, same as ruby

begin  
  raise 'bad'  
rescue  
  x = 1  
end  
x + 1  # => 2  

With js I need to define x ahead of time:

let x;  
try {  
  throw Error("Bad");  
} catch {  
  let x = 1;  
}  
x + 1;  // => 2  

Run pytest with verbose output

pytest -svv  

Set a breakpoint

In modern versions of python breakpoint() will suffice.
No need to import ‘pdb’ explicitly.

TODO: This file needs to be formatted.

Also see ~/notes/pip.md
Also see ~/notes/pyenv.md

python, inspiration, live coding, David Beazley, performance, concurrency, GIL, coroutines

https://www.youtube.com/watch?v=MCs5OvhV9S4
Source: https://news.ycombinator.com/item?id=36785005

python simple http server, SimpleHTTPServer, python 3

Serve current directory at http://localhost:8000
python -m http.server

python, repl, save file

After prototyping in the repl, dump all code I’ve entered into the repl:
%save my_file.py
So awesome!

python, debugging

Examine MRO with:
class Foo
pass
class Bar:
pass
class Baz(Foo, Bar):
pass
Baz.__mro__

python, gotchas

https://github.com/satwikkansal/wtfpython

pdb, which file am I in

p file

pdb, command list

https://docs.python.org/3.6/library/pdb.html#pdbcommand-step
unt is helpful to get out of a loop
Also use up and down

ipython turn off autocompletions

ipython --TerminalInteractiveShell.autosuggestions_provider=None  

ipython, mypy

Check mypy types in the repl with:
pip install mypy_ipython
ipython
%load_ext mypy_ipython
… do some prototyping
%mypy

pdb, pytest

run tests with –pdb to drop you into a repl on test failure

pytest, mocking

Use mock.mock_calls to get a list of all calls sent to a mock.
Or mock.call_args_list. I’m not sure what the difference is.

x = ‘hello’
print(f’{x = }’)

Also see alternatives ‘ice cream’ and ‘q’: https://github.com/zestyping/q

pdbpp

Use bt to print frames, then
f [number]
to jump to one

python modify path

PYTHONPATH=“./:$PYTHONPATH” python path/to/script.py

read unit.xml files

Install https://github.com/lukejpreston/xunit-viewer
Then
pytest –junit-xml=build/unit.xml
xunit-viewer -r build/unit.xml

python, pypi, list versions, pip list versions, hack

pip install my-package==

python zen of python

import this

python testing sibling library

cd ~/dev/my_app
source .env/bin/activate
pip install -e ~/dev/my_lib

python path, pythonpath, modify path, run from vim

:!PYTHONPATH=“app:$PYTHONPATH” python %

files, naming, modules

.py files can be imported as a module, but only if they do not contain hyphens in the name
bad: from bad-thing import foo
good: from good_thing import foo

bdist, sdist

See ~/dev/snippets/python/example_package/README.md and the associated
package layout

sdist is a source distribution, build with:
python setup.py sdist
Build wheel distribution with:
python setup.py sdist bdist_wheel

Note that .whl files can be unzipped with unzip

The distributions are saved in the relative dist/ directory. Untar the source
distribution or unzip the .whl distribution to debug.

To include a non-python file in the wheel, the setting include_package_data
must be True in setup.py AND an include statement must be present in
MANIFEST.in

Wheel is the more modern form of egg: https://packaging.python.org/discussions/wheel-vs-egg/
> Wheel is currently considered the standard for built and binary packaging for Python.

Good slides here: https://blog.ionelmc.ro/presentations/packaging/#slide:1

debugging, python, repl

Pickle is fully specified to reconstruct types
Insert breakpoint in problematic process, pickle.dump to file
pickle.load in script with debugging aids

mypy, readme

http://calpaterson.com/mypy-hints.html

python get size of object, memory, ram, size in bytes

import sys
sys.getsizeof(full_dataset)
See ~/notes/random_notes_worth_keeping.txt search ‘memory usage’

module not found

Seeing: ModuleNotFoundError: No module named ‘.’
Fix: Add an init.py file into

pip install git branch

pip install git+https://github.com//@

python and &, numpy, vectorized and

https://stackoverflow.com/questions/22646463/and-boolean-vs-bitwise-why-difference-in-behavior-with-lists-vs-nump/22647006#22647006

python2 to python3, migration, equivalence

http://diveinto.org/python3/porting-code-to-python-3-with-2to3.html#next

python mocking

Most articles I’ve found on python mocking have been junk. Collection of easy to follow tips:
https://wesmckinney.com/blog/spying-with-python-mocks/
https://stackoverflow.com/a/25424012/143447

environment variables, env variables, python, virtual env

find where package is located, installed

$ python -v
> import

python debugging, pdb

Install pdb++ https://pypi.org/project/pdbpp/
pip install pdbpp
Then at the debugger prompt, type ‘sticky’ for a much better debugging experience

python debugging, pdb, multiline

type ‘interact’ at the pdb prompt

virtual environment, virtual env, python2

$ /usr/local/Cellar/python@2/2.7.15_2/bin/pip2.7 install virtualenv
$ python2.7 /usr/local/lib/python2.7/site-packages/virtualenv.py .env_2_7
$ source .env_2_7/bin/activate

virtualenv, system python, python 2.7

sudo /usr/bin/easy_install-2.7 pip
/usr/bin/python2.7 -m pip install virtualenv
cd
/usr/bin/python2.7 -m virtualenv .env
source .env/bin/activate
pip install -r requirements.txt

pdb, list, display, source, return to breakpoint

list .

ipython enter newline

It’s tricky to add code to a previous code block after up-arrowing to it. To insert a newline, use:
ctlr+o+n or ctrl+q+j

ipython read file, ipython load, ipython load script

%load filename.py

ipython, pdb

Use pp to pretty print objects. Not built in anymore?

from pprint import pprint as pp  

pytest hides prints, show print, pytest

pytest -s to show print statements when you run tests
very confusing that the default setting eats print statements
pytest -s -vv to show verbose difference

pytest stop after first failure

pytest -x

pytest run single test

pytest -k ‘test_the_thing’

pytest run single file

pytest path/to/file.py

ipython reload, autoreload, repl

%load_ext autoreload
%autoreload 2
Or add this to ~/.ipython/profile_default/ipython_config.py to automatically
reload on every session:
c.InteractiveShellApp.exec_lines = [‘%load_ext autoreload’, ‘%autoreload 2’]

ipython profile, python profile

%prun some_function

virtual environment, python2

$ virtualenv .env

python create environment, python virtual environment, python3

python -m venv .env
(add env to gitignore)
source .env/bin/activate
[optional]
see ~/notes/jupyter.txt
deactivate

venv, virtual env, requirements, freeze

pip freeze > requirements.txt

python env, git

Add .env to .gitignoremy_environment
python -m venv .env
source .env/bin/activate
pip install -r requirements.txt

readme

https://access.redhat.com/blogs/766093/posts/2592591

ipython, matplotlib, plots, graph, math

ipython –matplotlib
See ~/dev/snippets/python/matplotlib_experiment.py for an example

Within the plot window, use control to pan (drag along the x or y axis), or use
the zoom rect feature on the toolbar

ipdb, single variable names

Say you have a variable ‘n’. Use exclamation point in front:
ipdb> !n

string match, methods matching, array select, list includes

Find all methods that have pid in it:
[x for x in dir(os) if re.search(‘pid’, x, re.I)] # => re.I for ignore case

getting help, help text, helptext, docstring

help(s.listen)
or
print(s.listen.__doc__)

from future import print_function

Getting list of methods:
> dir(StringIO.StringIO)

Empty class implementation:
class Foo:
pass

Debugging

  Get class name:  
  ipdb> x.__class__  
  or  
  ipdb> x.__class__.__name__  

Packages

  $ pip install numpy  
  $ pip install matplotlib  
  $ pip install ipdb   # => debugger  

Autocompletion:

              $ pip install jedi  
~/.vim/bundle $ git clone --recursive https://github.com/davidhalter/jedi-vim.git  

Interactive python:

$ python -i  
Better:  
$ ipython  

System python installs to
pip installs things to: /usr/local/lib/python2.7/site-packages

Pretty printing:

from pprint import pprint as pp  

Longer pretty printing:

import pprint  
params = {  
    'latitude': 37.775818,  
    'longitude': -122.418028,  
    'server_token': 'snip'  
}  
pp = pprint.PrettyPrinter(indent=4, width=1).pprint  
pp(params)  

Pretty printing something that resembles a dictionary

pp(dict(response.headers))   # => note the width of 1 above is important  

Patches
### python debugger, ipdb, list, patch
In /home/lou/dragon/lib/python2.7/site-packages/IPython/core/debugger.py,
Changed context from 3 to 10:

def print_stack_entry(self,frame_lineno,prompt_prefix='\n-> ',  
                      context = 10):  

Configuring iPython

Create default profile:  
$ ipython profile create  
$ ipython locate profile  

iPython tricks

Run a script:  
> %run workbench.py  

Insert an enter (newline) in repl:

ctrl+v + ctlr+j  

Also try %edit

Exec from command line

python -c "print('sup')"  

Sys

import sys  
sys.path  

Entry detection

if __name__ == '__main__':  
    print("yes, I am main")  

Trick to assign characters to fixed 0 to 25 slots

    slots = [0] * 26  
    c = 'x'  
    slot = ord(c) - ord('a')  
    slots[slot] += 1  

Opposite of ord

    chr(97) # => 'a'  

How to apply typing to generators, genexp

The arguments are:

Using a python generator as a coroutine (before async def)

from typing import Generator  

def fn() -> Generator[str, str, None]:  
    name: str = yield "I am now primed"  
    yield f"{name} world"  

g = fn()  

# Prime the generator. The argument needs to be None, otherwise this will raise:  
# TypeError: can't send non-None value to a just-started generator"  
print(g.send(None))  

# Now I can get the coroutine to spit out "hello world"  
print(g.send("hello"))  

The above is much better written as:

import asyncio  

async def fn(name: str) -> str:  
    return f"{name} world"  

async def main():  
    result = await fn("hello")  
    print(result)  

asyncio.run(main())  

Python gotcha, no implicit self

Bad:

class C:  
    def __init__(self):  
        _x = "x"  

Good:

class C:  
    def __init__(self):  
        self._x = "x"  

Python get context switch interval

sys.getswitchinterval()  

Change it with:

sys.setswitchinterval(1E-2)  

Python delegate pattern

Not sure how many python devs use this, but a translation of obj-c/swift delegate pattern could be:

class C:  
    def __init__(self, delegate):  
        self.delegate = delegate  
    def dowork(self):  
        self.delegate.did_finish()  

class D:  
    def did_finish(self):  
        print("work is done")  

c = C(D())  
c.dowork() # => "work is done"  

Or, with typing:

from typing import Protocol  

class P(Protocol):  
    def did_finish(self) -> None:  
        ...  

class C:  
    def __init__(self, delegate: P):  
        self.delegate = delegate  
    def dowork(self) -> None:  
        self.delegate.did_finish()  

class D(P):  
    def did_finish(self) -> None:  
        print("work is done")  

c = C(D())  
c.dowork()  

Python callback pattern

Using callbacks with pluggable closures

from collections.abc import Callable  

class C:  
    def __init__(self, did_finish: Callable[[], None]):  
        self.did_finish = did_finish  
    def dowork(self) -> None:  
        self.did_finish()  

def done() -> None:  
    print("work is done")  

c = C(done)  
c.dowork()  

Or

c = C(lambda: print("work is done"))  
c.dowork()  

Or

class X:  
    def __call__(self) -> None:  
        print("work is done")  

c = C(X())  
c.dowork()  

Cartesian product

from itertools import product  
list(product(('a', 'b'), (1, 2)))  
# => [('a', 1), ('a', 2), ('b', 1), ('b', 2)]  

Or

list(product(('a', 'b'), repeat=2))  
[('a', 'a'), ('a', 'b'), ('b', 'a'), ('b', 'b')]  

Random numbers

Ints

from random import randint  
randint(0, 10) # => element in [0, 10]  

Or

from random import randrange  
randrange(0, 10) # => element in (0, 10]  

Or

from random import choice  
choice(range(10))  

Use listcomps instead of maps

Say I want a list of random numbers in [1, 5]. I naturally think:

map(lambda _: randint(1,5), range(10))  

but it’s more idiomatic to:

[randint(1,5) for _ in range(10)]  

Or, better:

from random import choices  
choices(range(1, 6), k=10)  

Or, by far the fastest for large N:

import numpy as np  
np.random.randint(0, 11, size=100000)  

Python gotcha, mistaken use of class variable

Example from: https://docs.python.org/3/tutorial/classes.html

class Dog:  
    tricks = []             # mistaken use of a class variable  

    def __init__(self, name):  
        self.name = name  

    def add_trick(self, trick):  
        self.tricks.append(trick)  

d = Dog('Fido')  
e = Dog('Buddy')  
d.add_trick('roll over')  
e.add_trick('play dead')  
d.tricks  
['roll over', 'play dead']  

Python gotcha, pop(0) is O(n)

[1,2,3].pop()  # => 3 in O(1)  
[1,2,3].pop(0) # => 1 in O(n)  

Use deque instead (double ended queue):

from collections import deque  
d = deque([1,2,3])  
d.pop()     # => 3 in O(1)  
d.popleft() # => 1 in O(1)  

deque also has automatic eviction:

d = deque([1,2,3], maxlen=3)  
d.append(4)  
d  
# => [2,3,4]  

canonical BFS with a deque

Also remember for a known-good tree, the visited set is not necessary:

visited = {start}  
queue = deque([start])  

while queue:  
    node = queue.popleft()  

    for neighbor in graph[node]:  
        if neighbor not in visited:  
            visited.add(neighbor)  
            queue.append(neighbor)  

Python gotcha, methods that mutate return None

x = [1,2].append(3)  
x # => none  

Python insort

from bisect import bisect, bisect_left  

x = [0,1,1,2]  
bisect(x, 1) # => 3  
bisect_left(x, 1) # => 1  

from bisect import insort  
insert(x, 1.5)  
x # => [0,1,1,1.5,2]  

Get binary of a base 10 number

 bin(255)  
 # => '0b11111111'  

How to create a generic type (for mypy)

from typing import TypeVar  
from typing import Generic  

T = TypeVar('T')  

class Node(Generic[T]):  
    def __init__(self, value: T):  
        self.value = value  

node1: Node[int] = Node('1')  # This will throw a mypy error  
node2: Node[int] = Node(1)    # mypy happy  

How to create a generic fn

I don’t need to use Generic for this. It’s just:

from typing import TypeVar  
T = TypeVar("T")  
def rotate(list: list[T]) -> list[T]:  
    ...  

How to encode and decode to utf8

"hello".encode('utf-8')  
bin.decode('utf-8')  

And ‘utf-8’ is the default, so I can actually use:

"hello".encode()  
bin.decode()  
  

How to read a file in chunks

In this example I’m using rb and then decoding as utf-8.
It’s not necessary to do this, but I want to remember how to read as binary:

with open('read_in_chunks.txt', 'rb') as f:  
while binary := f.read(2):  
    print(binary.decode('utf-8'))  

How to read from an IO stream in chunks

from io import BytesIO  

with BytesIO(b'hello world') as stream:  
    print(stream.read(3))  

Range in descending order

Note that just like an ascending range, the range is not inclusive of the second argument:

list(range(3, 0, -1))  # => [3, 2, 1]  
  

Check truthiness, truthy

Use string formatters

a = f'{15:x}'                # => 'f'  
b = format(15, 'x')          # => 'f'  
c = '{:x}'.format(15)        # => 'f'  
d = '{num:x}'.format(num=15) # => 'f'  

Benefit of __slots__

More efficient, sure. But what I really like is this:

...with the default __dict__, a misspelled variable name results in the  
creation of a new variable, but with __slots__ it raises in an  
AttributeError.   

Source: https://wiki.python.org/moin/UsingSlots  

For example:

class C:  
    def __init__(self):  
        self._x = None  

    def run(self):  
        self.x = 1  

c = C()  
c.run() # => Nothing raised  
c.__dict__ # => contains `_x` and `x`  

versus:

class C:  
    __slots__ = ('_x',)  

    def __init__(self):  
        self._x = None  

    def run(self):  
        self.x = 1  

c = C()  
c.run() # => Raises "AttributeError: 'C' object has no attribute 'x'"  

Getters and setters

The setter syntax is a little funky, and relies on the existance of the getter.

 class C:  
     @property  
     def p(self):  
         return self._p  

     @p.setter  
     def p(self, value):  
         self._p = value  

Frozenset usage

frozenset is helpful to make things hashable, but you can’t use it blindly.
Say I was trying to use [‘a’, ‘b’] as a key to a dict. I can’t do this:

{['a', 'b']: 'x'}  # => Raises "TypeError: unhashable type: 'list'"  

but I could do this:

fs1 = frozenset(['a', 'b'])  
d = {fs1: 'x'}  

but beware (note the ordering of the array):

fs2 = frozenset(['b', 'a'])  
d[fs2] # => 'x'  

Use a tuple if order and duplicates matter as the key. Tuples are hashable (unlike swift!)

How to put a set inside a set:

Wrong:

set(frozenset())   # => Not what I expect, but it makes sense. `frozenset()` is an iterable that the `set(...)` initializer iterates over to set the initial contents of the set.  

Right:

set([frozenset()])  

Flatten

There is nothing like ruby’s built in [1, [2,3]].flatten # => [1,2,3].

For a list of lists, from_iterable works well:

from itertools chain  
list(chain.from_iterable([[1], [2, 3]])) # => [1,2,3]  

but for [1, [2, 3] example, I’d have to use type inspection:

from itertools import chain  
l = [1,[2,3]]  
chain.from_iterable(x if isinstance(x, list) else [x] for x in l)  

for well structured, can also just use:

inp = [[1], [2, 3]]  
[x for item in inp for x in item]  

remember the trick to these nested list comprehensions is that the outer loop is the first for, e.g. for item in inp

Regex extracting groups

x = re.match(r'a(b)', 'ab')  
x.group(0)  
# => 'ab'  
x.group(1)  
# => 'b'  

Counters can take a sequence as an initializer arg

from collections import Counter  
Counter('xyzz')['z'] #=> 2  

Counters support + overload

c = Counter()  
c['a'] += 1  
d = Counter()  
d['a'] += 1  
d['b'] = 99  
c + d  
# => Counter({'b': 99, 'a': 2})  

Matrix traversal, iterate diagonal

Two facts that make all this simpler:

  1. r - c = d for any diagonal d that I’m iterating

  2. There are m + n - 1 diagonals

     matrix = [  
         [1, 3, 5, 7],  
         [2, 4, 6, 8],  
         [9, 11, 13, 15]  
     ]  
    
     m = len(matrix)  
     n = len(matrix[0])  
    
     print("Iterate all diagonals:")  
     for d in range(-n + 1, m):  
         for r in range(m):  
             c = r - d  
             if 0 <= c < n:  
                 print(f'visiting {matrix[r][c]=}')  
    
    
     print("Iterate diagonals that start at the top row:")  
     for d in range(-1, -n, -1):  
         for c in range(n):  
             r = c + d  
             if 0 <= r < m:  
                 print(f'visiting {matrix[r][c]=}')  

Create an mxn matrix

 m = 5  
 n = 3  
 nums = list(range(1, n * m + 1))  
 matrix = [nums[r*n:(r+1)*n] for r in range(m)]  

Handy trick to get the matrix by columns

Amazing that this transposes a matrix:

list(zip(*matrix))  

If I need to transpose in place (for a square matrix):

n = len(matrix)  
for r in range(n):  
    for c in range(r + 1, n):  
        matrix[r][c], matrix[c][r] = matrix[c][r], matrix[r][c]  

Nice trick to get the max k/v out of a dictionary

counts = {"a": 3, "b": 7, "c": 2}  
max(counts, key=counts.get)  # => "b"  

If I need the max k, a counter is handy:

from collections import Counter  
counts = Counter({"a": 3, "b": 7, "c": 2})  
counts.most_common(k)  

Threading reference, and a gotcha

If I need a mutex around some shared state, use:

from threading import Lock  
unfair_lock = Lock()  
with unfair_lock:  
    # my critical section  

To submit work to a pool of background threads, use:

from concurrent.futures import ThreadPoolExecutor  
with ThreadPoolExecutor() as pool:  
    for res in pool.map(add_one, [1,2,3]):  
        print(  
    # Submit work  

The API is very thin for ThreadPoolExecutor, just submit, map, and shutdown.

If I submit work with pool.submit, I get back a future. Work is scheduled immediately and does not block.

Gotchas:
- future.result blocks until that tasks’s result is ready
- pool.map(fn, argument) returns an in-order iterator of results. Calling next or iterating in a for blocks until that task’s result is ready.
- The main thread becomes the coordination thread, and it will block (!) until all work is done in the contextmanager, even if I don’t explicitly wait on results.

Heap reference

Example straight from the heap docs: https://docs.python.org/3/library/heapq.html#basic-examples

from heapq import heappush, heappop  

h = []  
heappush(h, (5, 'write code'))  
heappush(h, (7, 'release product'))  
heappush(h, (1, 'write spec'))  
heappush(h, (3, 'create tests'))  
heappop(h)  

Another good one is heapify:

heap = [5, 2, 8]  
heapq.heapify(heap)  
heapq.heappush(heap, 1)  
heapq.heappop(heap)      # 1  
heap[0]                  # 2  

Ref:

Operation Meaning Time
heapq.heapify(nums) Turn a list into a heap, in place O(n)
heapq.heappush(heap, x) Add an item O(log n)
heapq.heappop(heap) Remove and return the smallest O(log n)
heap[0] Peek at the smallest O(1)
heapq.heappushpop(heap, x) Push, then pop the smallest O(log n)
heapq.heapreplace(heap, x) Pop the smallest, then push O(log n)

OrderedDict reference

OrderedDict gives you move_to_end and popitem(last=False). Can use it for an easy LRU cache.

No need to manage a dict and doubly linked list myself.

Also, somewhere I got it in my head that I needed to evict before adding,
but I don’t. Add then evict makes the logic straightforward:

from collections import OrderedDict  

class LRUCache:  
    def __init__(self, capacity):  
        self.capacity = capacity  
        self.od = OrderedDict()  

    def get(self, key):  
        if key not in self.od:  
            return -1  

        self.od.move_to_end(key)  
        return self.od[key]  

    def put(self, key, val):  
        self.od[key] = val  
        self.od.move_to_end(key)  

        if len(self.od) > self.capacity:  
            self.od.popitem(last=False)  

Update: I know where evict before adding came from, that’s an LFU cache, because
if you insert first that k/v’s used counter is at one and would be immediately evicted.

hashlib reference

Handy for creating hashes of arbitrary files:

import hashlib  

hasher = hashlib.sha256()  
with open(path, "rb") as f:  
    while chunk := f.read(64 * 1024):  
        hasher.update(chunk)  
result = hasher.hexdigest()  

Git uses SHA-1. There is also a direct form if I don’t need to stream in chunks:

hashlib.sha256(b'hello world')  

Massive footgun in asyncio

If you forget to call to_thread on a sync function, you block the entire event loop!

Threading and concurrency

It’s a big topic: https://docs.python.org/3/library/concurrency.html

Deque reference

Most used are the initializer with maxlen for built-in eviction, append, and popleft

and : deque(maxlen=3)

Operation Code Returns / effect Time
Initialize w eviction deque(maxlen=3) new deque O(1)
Initialize w existing deque([1,2,3]) new deque O(n)
Add right q.append(4) [1, 2, 3, 4] O(1)
Add left q.appendleft(0) [0, 1, 2, 3] O(1)
Remove right q.pop() Returns 3 O(1)
Remove left q.popleft() Returns 1 O(1)
Peek right q[-1] 3 O(1)
Peek left q[0] 1 O(1)
Length len(q) 3 O(1)
Empty check not q False O(1)
Extend right q.extend([4, 5]) [1, 2, 3, 4, 5] O(k)
Extend left q.extendleft([4, 5]) [5, 4, 1, 2, 3] — reverses input O(k)
Rotate right q.rotate(1) [3, 1, 2] O(k) for k steps
Rotate left q.rotate(-1) [2, 3, 1] O(k) for k steps
Reverse in place q.reverse() [3, 2, 1] O(n)
Membership 2 in q True O(n)
Count q.count(2) 1 O(n)
Remove first match q.remove(2) [1, 3]; raises ValueError if absent O(n)
Clear q.clear() Empty deque O(n)
Find index q.index(val) Returns the first index of the value O(n)
Insert at q.insert(index, val) Inserts value before index O(n)